Large Language Models as Clinical Decision Support for Diagnosis: Benefit After AI-Literacy Training Versus Hallucination and Automation Bias
This narrative review examines the conditions under which physicians benefit from large language models (LLMs) as clinical decision support systems (CDSS) for diagnosis, with particular focus on the moderating role of structured AI-literacy training and the persistent safety risks of hallucination and automation bias. BACKGROUND: Diagnostic error remains a leading source of preventable patient harm globally. LLMs such as GPT-4, Med-PaLM 2, and specialized clinical models have demonstrated strong performance on standardized medical examinations, generating interest in their deployment as diagnostic assistants. However, LLMs are generative, probabilistic systems that differ fundamentally from traditional rule-based CDSS, introducing new categories of clinical risk. METHODS: This review synthesizes evidence from randomized controlled trials and observational studies conducted across diverse healthcare settings, including the United States, the Netherlands, Indonesia, and Kenya. Key trials examined include a US-based RCT evaluating GPT-4-assisted diagnosis (n=50), resource-limited setting trials from Indonesia and Kenya, and a landmark NEJM AI automation-bias trial involving a 20-hour AI-literacy training intervention. KEY FINDINGS: (1) In developed healthcare systems, LLM access without structured training produced no significant diagnostic improvement (US: 76% vs. 74%, p=0.60). (2) In resource-limited settings, LLM access significantly improved diagnostic accuracy (+7.9% to +15.1%, p<0.001), effectively substituting for unavailable specialist consultation and limited diagnostic infrastructure. (3) Even after comprehensive 20-hour AI-literacy training, physicians remained highly vulnerable to automation bias, accepting erroneous LLM recommendations at a rate of 84.9%. (4) Three mechanisms mediate LLM benefit: better prompting, limitation awareness, and systematic verification behavior. PROPOSED FRAMEWORK: The evidence supports a layered safety model combining AI-literacy training (Layer 1) with a structured verification workflow (Layer 2). The four-step workflow requires clinicians to: (1) Form an independent clinical assessment before viewing LLM output; (2) Request LLM output as a supervised second reader; (3) Verify recommendations against patient history, guidelines, labs, and imaging; (4) Document the decision to accept, modify, or reject. The core principle is "Trained Clinician + Verified LLM Workflow" — never "LLM Alone." REGULATORY CONTEXT: The proposed framework aligns with the EU AI Act and Medical Device Regulation (MDR 2017/745), which classify diagnostic LLMs as high-risk AI systems requiring human oversight, conformity assessment, and post-market surveillance. CONCLUSION: LLMs can be valuable diagnostic CDSS, particularly in resource-constrained environments. However, hallucination and automation bias remain major safety risks that training alone cannot eliminate. The safest path forward requires both AI-literate clinicians and engineered workflow safeguards. This publication includes the full review article (10 pages, 6 figures) and supplementary data visualizations.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex