The promise of artificial intelligence (AI) to improve and reduce inequities in access, quality, and appropriateness of high-quality diagnosis remains largely unfulfilled. Vast clinical data sets, extensive computational capacity, and highly developed and accessible machinelearning tools have resulted in numerous publications that describe high-performing algorithmic approaches for a variety of diagnostic tasks. However, such approaches remain largely unadopted in clinical practice. This discrepancy between promise and practice— the AI chasm—has many causes. Some reasons are endemic to the larger field of AI, including a lack of generalizability and reproducibility for the published algorithms. Other reasons are more specific to clinical AI, such as a lack of gender, racial, and ethnic diversity in clinical data sets and insufficient evaluation of the algorithms in clinical settings. The disconnect between the metrics for algorithm performance and the realities of a clinician’s workflow and decision-making process is a fundamental but often overlooked issue. The inclusion of clinical context in AI performance metrics for optimizing and evaluating clinical algorithms could make AI tools more clinically relevant and readily adopted (Box).
Paper
Full text
Rethinking Algorithm Performance Metrics for Artificial Intelligence in Diagnostic Medicine
Semantic Scholar · Medicine · 2022
Abstract
The promise of artificial intelligence (AI) to improve and reduce inequities in access, quality, and appropriateness of high-quality diagnosis remains largely unfulfilled. Vast clinical data sets, extensive computational capacity, and highly developed and accessible machinelearning tools have resulted in numerous publications that describe high-performing algorithmic approaches for a variety of diagnostic tasks. However, such approaches remain largely unadopted in clinical practice. This discrepancy between promise and practice— the AI chasm—has many causes. Some reasons are endemic to the larger field of AI, including a lack of generalizability and reproducibility for the published algorithms. Other reasons are more specific to clinical AI, such as a lack of gender, racial, and ethnic diversity in clinical data sets and insufficient evaluation of the algorithms in clinical settings. The disconnect between the metrics for algorithm performance and the realities of a clinician’s workflow and decision-making process is a fundamental but often overlooked issue. The inclusion of clinical context in AI performance metrics for optimizing and evaluating clinical algorithms could make AI tools more clinically relevant and readily adopted (Box).