The integration of artificial intelligence (AI) into healthcare presents an unprecedented opportunity to improve access to quality care, particularly in fields facing workforce shortages and diagnostic complexity. Dermatology exemplifies this challenge, with stark disparities in specialist access both within high-income countries and globally. These inequities are especially pronounced for individuals with rare skin disorders, such as genodermatoses—genetic conditions with multisystem involvement that are frequently underdiagnosed, misdiagnosed, and undertreated. Large language models (LLMs), such as ChatGPT, have demonstrated growing capacity to generate medically relevant content and synthesize complex information rapidly. Their potential applications in rare diseases are promising, including support in diagnosis, patient education, and clinical decision-making. However, the utility of LLMs in dermatology, particularly for rare genetic conditions, remains underexplored. This paper highlights the intersection of AI, health equity, and rare disease care, emphasizing the urgent need for robust evaluation frameworks to assess LLM outputs for accuracy, safety, clarity, empathy, and trustworthiness. As global health efforts aim to meet the needs of millions living with rare diseases, rigorous validation of AI tools is critical to ensure they reduce, rather than reinforce, existing disparities.
Paper
Full text
Evaluating the performance of large language models in diagnosing rare genodermatoses
Semantic Scholar · 2025
Abstract
The integration of artificial intelligence (AI) into healthcare presents an unprecedented opportunity to improve access to quality care, particularly in fields facing workforce shortages and diagnostic complexity. Dermatology exemplifies this challenge, with stark disparities in specialist access both within high-income countries and globally. These inequities are especially pronounced for individuals with rare skin disorders, such as genodermatoses—genetic conditions with multisystem involvement that are frequently underdiagnosed, misdiagnosed, and undertreated. Large language models (LLMs), such as ChatGPT, have demonstrated growing capacity to generate medically relevant content and synthesize complex information rapidly. Their potential applications in rare diseases are promising, including support in diagnosis, patient education, and clinical decision-making. However, the utility of LLMs in dermatology, particularly for rare genetic conditions, remains underexplored. This paper highlights the intersection of AI, health equity, and rare disease care, emphasizing the urgent need for robust evaluation frameworks to assess LLM outputs for accuracy, safety, clarity, empathy, and trustworthiness. As global health efforts aim to meet the needs of millions living with rare diseases, rigorous validation of AI tools is critical to ensure they reduce, rather than reinforce, existing disparities.