Retrieval-augmented generation for medication safety: A case study using drug package inserts.

BACKGROUND Retrieval-augmented generation (RAG) has shown promise in mitigating hallucinations in large language models (LLMs), although its utility in medication safety education remains underexplored. To evaluate the performance of RAG-based LLMs in generating patient medication safety education materials, and to assess the agreement between LLM-as-a-judge evaluations and clinical expert judgments. METHODS Two senior clinical pharmacists selected medications and defined key safety education entries and developed an ontology mapping these entries to sections of drug package inserts. Using this ontology, GPT-4o, OAGLLM, and the pharmacists each independently generated educational materials. Seven licensed clinical pharmacists conducted blinded evaluations using a 5-point Likert scale across six dimensions (Turing Test, Coverage, Relevance, Accuracy, Harmfulness, and Severity of Harm), with each entry independently assessed by one domain-matched pharmacist. DeepSeek-V3, DeepSeek-R1, Qwen-Plus, Claude Sonnet 4.5 also assessed the materials using the same evaluation criteria as the judge models. Inter-rater agreement was assessed using weighted kappa and Spearman correlation coefficients. RESULTS Across 70 medications, human experts achieved the highest overall score (0.89 ± 0.16), followed by GPT-4o (0.86 ± 0.20) and OAGLLM (0.84 ± 0.17). Human-generated materials performed best in Coverage, Relevance, and Accuracy, while GPT-4o achieved the lowest Harmfulness score. Evaluation results showed that the four judged-LLMs achieved fair agreement (κ = 0.21-0.24) and low-to-moderate alignment (ρ = 0.40-0.53) with human experts. CONCLUSIONS While clinical experts remain superior in generating medication safety education, GPT-4o demonstrates encouraging potential. However, the limited agreement between LLM-based evaluations and human judgments highlights the ongoing need for expert oversight.

Paper

Full text

PDF

Retrieval-augmented generation for medication safety: A case study using drug package inserts.

Semantic Scholar · Medicine · 2026

Abstract

BACKGROUND Retrieval-augmented generation (RAG) has shown promise in mitigating hallucinations in large language models (LLMs), although its utility in medication safety education remains underexplored. To evaluate the performance of RAG-based LLMs in generating patient medication safety education materials, and to assess the agreement between LLM-as-a-judge evaluations and clinical expert judgments.

METHODS Two senior clinical pharmacists selected medications and defined key safety education entries and developed an ontology mapping these entries to sections of drug package inserts. Using this ontology, GPT-4o, OAGLLM, and the pharmacists each independently generated educational materials. Seven licensed clinical pharmacists conducted blinded evaluations using a 5-point Likert scale across six dimensions (Turing Test, Coverage, Relevance, Accuracy, Harmfulness, and Severity of Harm), with each entry independently assessed by one domain-matched pharmacist. DeepSeek-V3, DeepSeek-R1, Qwen-Plus, Claude Sonnet 4.5 also assessed the materials using the same evaluation criteria as the judge models. Inter-rater agreement was assessed using weighted kappa and Spearman correlation coefficients.

RESULTS Across 70 medications, human experts achieved the highest overall score (0.89 ± 0.16), followed by GPT-4o (0.86 ± 0.20) and OAGLLM (0.84 ± 0.17). Human-generated materials performed best in Coverage, Relevance, and Accuracy, while GPT-4o achieved the lowest Harmfulness score. Evaluation results showed that the four judged-LLMs achieved fair agreement (κ = 0.21-0.24) and low-to-moderate alignment (ρ = 0.40-0.53) with human experts.

CONCLUSIONS While clinical experts remain superior in generating medication safety education, GPT-4o demonstrates encouraging potential. However, the limited agreement between LLM-based evaluations and human judgments highlights the ongoing need for expert oversight.

Similar papers

© 2026 NYSGPT2525 LLC