Related Experiment Video
Updated: Jul 13, 2026

Drug Repurposing Hypothesis Generation Using the "RE:fine Drugs" System
Published on: December 11, 2016
Retrieval-augmented generation for medication safety: A case study using drug package inserts
Yuanyuan Sun1, Wei Wang2, Wanqiu Cheng3
1Center for Clinical Pharmacology, The Third Xiangya Hospital, Central South University, Changsha 410013, China.
Background:
Retrieval-augmented generation (RAG) has shown promise in mitigating hallucinations in large language models (LLMs), although its utility in medication safety education remains underexplored. To evaluate the performance of RAG-based LLMs in generating patient medication safety education materials, and to assess the agreement between LLM-as-a-judge evaluations and clinical expert judgments.
Methods:
Two senior clinical pharmacists selected medications and defined key safety education entries and developed an ontology mapping these entries to sections of drug package inserts. Using this ontology, GPT-4o, OAGLLM, and the pharmacists each independently generated educational materials. Seven licensed clinical pharmacists conducted blinded evaluations using a 5-point Likert scale across six dimensions (Turing Test, Coverage, Relevance, Accuracy, Harmfulness, and Severity of Harm), with each entry independently assessed by one domain-matched pharmacist. DeepSeek-V3, DeepSeek-R1, Qwen-Plus, Claude Sonnet 4.5 also assessed the materials using the same evaluation criteria as the judge models. Inter-rater agreement was assessed using weighted kappa and Spearman correlation coefficients.
Results:
Across 70 medications, human experts achieved the highest overall score (0.89 ± 0.16), followed by GPT-4o (0.86 ± 0.20) and OAGLLM (0.84 ± 0.17). Human-generated materials performed best in Coverage, Relevance, and Accuracy, while GPT-4o achieved the lowest Harmfulness score. Evaluation results showed that the four judged-LLMs achieved fair agreement (κ = 0.21-0.24) and low-to-moderate alignment (ρ = 0.40-0.53) with human experts.
Conclusions:
While clinical experts remain superior in generating medication safety education, GPT-4o demonstrates encouraging potential. However, the limited agreement between LLM-based evaluations and human judgments highlights the ongoing need for expert oversight.
More Related Videos
Related Concept Videos
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Dosage Regimens: Designs and Approaches
Pharmaceutical Poisoning: Potential Scenarios
Drug Dosing: Geriatric Patients
Drug Toxicity: Risk factors
Dosage Regimen: Individualization

