RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
Numerous medical systems powered by Large Language Models (LLMs) have achieved substantial progress, enabling them to perform diverse healthcare tasks. However, existing research is limited by the absence of real-world datasets specifically addressing medication safety, primarily due to privacy regulations and data accessibility challenges. Moreover, the evaluation of LLM-based systems in realistic clinical consultation settings, particularly with respect to medication safety, remains underexplored. To bridge these gaps, we propose a novel frame-work to simulate and evaluate clinical consultation scenarios for systematically assessing the medication safety capabilities of LLMs. Within this framework, we generate inquiry-diagnosis dialogues embedded with relevant medication risks and construct a dedicated medication safety database, RxRisk DB, comprising 6,725 contraindications, 28,781 drug interactions, and 14,906 indication-drug pairs. A two-stage filtering strategy ensures clinical realism and professional quality, resulting in the final benchmark, RxSafeBench, which includes 2,443 high-quality consultation scenarios evenly split across contraindication and interaction types. We evaluate state-of-the-art open-source and proprietary LLMs using a structured multiple-choice format that tests the models' ability to recommend the most appropriate medication given simulated patient context. Results reveal that current LLMs struggle to reliably incorporate contraindication and drug interaction information, especially when risks are implied rather than explicitly stated. Our findings highlight key challenges in deploying LLMs for medication safety and offer insights into improving their reliability through enhanced prompting strategies and task-specific fine-tuning. By introducing RxSafeBench, we provide the first comprehensive benchmark for assessing medication safety in LLMs, paving the way toward safer and more trustworthy AI-driven clinical decision support systems. Our code and data are released at https://github.com/CAS-SIAT-XinHai/RxSafeBench.
Paper
References (25)
Scroll for more · 13 remaining