Recent foundation models such as MedSAM, SAM-Med2D, and BiomedParse have shown promising results in medical image segmentation including skin lesion images. However, most studies focus on overall performance without examining how these models behave across clin-ically or visually distinct subgroups. In this study, we conduct a comprehensive evaluation of three foundation segmentation models on 10,015 dermoscopic images from the HAM10000 dataset. Beyond overall Dice scores, we compare model performance across annotated clinical at-tributes-including diagnosis, sex, age, and anatomical site-as well as fine-grained visual characteristics (artifact, lesion color, dermoscopic feature) grouped using MONET, a vision-language retrieval model. By leveraging MONET's ability to retrieve semantically similar images using natural language prompts (e.g., “hair,” “blue-white veil,” “pigment network”), we automatically constructed attribute-driven subgroups without manual labeling. Our results reveal con-sistent performance disparities across both annotated and MONET-derived subgroups, especially in underrepresented anatomical sites (ear and face), older patients, lesion type (basal cell carcinoma), and lesion with visualfeatures (der-moscope dark corner, yellow, blue, ulceration, blue-white veils). This is the first study to integrate MONET-based re-trieval for subgroup evaluation in lesion segmentation, pro-viding new insights into the robustness and trustworthiness of foundation models in dermatology.
Paper
Full text
Evaluating the Trustworthiness of Foundation Models for Skin Lesion Segmentation
Semantic Scholar · Computer Science · 2025
Abstract
Recent foundation models such as MedSAM, SAM-Med2D, and BiomedParse have shown promising results in medical image segmentation including skin lesion images. However, most studies focus on overall performance without examining how these models behave across clin-ically or visually distinct subgroups. In this study, we conduct a comprehensive evaluation of three foundation segmentation models on 10,015 dermoscopic images from the HAM10000 dataset. Beyond overall Dice scores, we compare model performance across annotated clinical at-tributes-including diagnosis, sex, age, and anatomical site-as well as fine-grained visual characteristics (artifact, lesion color, dermoscopic feature) grouped using MONET, a vision-language retrieval model. By leveraging MONET's ability to retrieve semantically similar images using natural language prompts (e.g., “hair,” “blue-white veil,” “pigment network”), we automatically constructed attribute-driven subgroups without manual labeling. Our results reveal con-sistent performance disparities across both annotated and MONET-derived subgroups, especially in underrepresented anatomical sites (ear and face), older patients, lesion type (basal cell carcinoma), and lesion with visualfeatures (der-moscope dark corner, yellow, blue, ulceration, blue-white veils). This is the first study to integrate MONET-based re-trieval for subgroup evaluation in lesion segmentation, pro-viding new insights into the robustness and trustworthiness of foundation models in dermatology.