D-LEMA: Deep Learning Ensembles from Multiple Annotations -- Application to Skin Lesion Segmentation
Medical image segmentation annotations suffer from inter- and intra-observer\nvariations even among experts due to intrinsic differences in human annotators\nand ambiguous boundaries. Leveraging a collection of annotators' opinions for\nan image is an interesting way of estimating a gold standard. Although training\ndeep models in a supervised setting with a single annotation per image has been\nextensively studied, generalizing their training to work with datasets\ncontaining multiple annotations per image remains a fairly unexplored problem.\nIn this paper, we propose an approach to handle annotators' disagreements when\ntraining a deep model. To this end, we propose an ensemble of Bayesian fully\nconvolutional networks (FCNs) for the segmentation task by considering two\nmajor factors in the aggregation of multiple ground truth annotations: (1)\nhandling contradictory annotations in the training data originating from\ninter-annotator disagreements and (2) improving confidence calibration through\nthe fusion of base models' predictions. We demonstrate the superior performance\nof our approach on the ISIC Archive and explore the generalization performance\nof our proposed method by cross-dataset evaluation on the PH2 and DermoFit\ndatasets.\n