Lizard: A Large-Scale Dataset for Colonic Nuclear Instance Segmentation and Classification

The development of deep segmentation models for computational pathology\n(CPath) can help foster the investigation of interpretable morphological\nbiomarkers. Yet, there is a major bottleneck in the success of such approaches\nbecause supervised deep learning models require an abundance of accurately\nlabelled data. This issue is exacerbated in the field of CPath because the\ngeneration of detailed annotations usually demands the input of a pathologist\nto be able to distinguish between different tissue constructs and nuclei.\nManually labelling nuclei may not be a feasible approach for collecting\nlarge-scale annotated datasets, especially when a single image region can\ncontain thousands of different cells. However, solely relying on automatic\ngeneration of annotations will limit the accuracy and reliability of ground\ntruth. Therefore, to help overcome the above challenges, we propose a\nmulti-stage annotation pipeline to enable the collection of large-scale\ndatasets for histology image analysis, with pathologist-in-the-loop refinement\nsteps. Using this pipeline, we generate the largest known nuclear instance\nsegmentation and classification dataset, containing nearly half a million\nlabelled nuclei in H&E stained colon tissue. We have released the dataset and\nencourage the research community to utilise it to drive forward the development\nof downstream cell-based models in CPath.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC