Related Experiment Video
Updated: Jul 17, 2026

Drug Repurposing Hypothesis Generation Using the "RE:fine Drugs" System
Published on: December 11, 2016
Real-world evaluation of medication recommendation workflows: Retrieval augmentation, physician-RAG collaborative
Long Ren1, Fang Fang1, Yuan Zhang1
1Department of Pharmacy, Shanghai East Hospital, Tongji University School of Medicine, Shanghai 200120,China.
Background:
Large language models (LLMs) are increasingly explored for clinical decision support, yet their role in medication recommendation remains uncertain. Prescribing requires not only selecting appropriate therapies, but also prioritizing essential treatments, avoiding unsafe medications, and specifying correct prescription details. Existing evaluations often emphasize benchmark reasoning rather than prescribing readiness in real-world settings.
Objective:
To compare four medication recommendation workflows using a clinically grounded framework focused on completeness, safety, ranking quality, and prescription-parameter accuracy.
Methods:
We conducted a retrospective real-world evaluation of 800 clinical cases from Shanghai East Hospital. A case-specific expert reference standard was constructed using three mutually exclusive categories: CORE, essential therapies; ALT, acceptable alternatives; and AVOID, unsafe or inappropriate medications. Four workflows were compared: Physician, Base LLM, Retrieval-augmented LLM, and Physician-RAG collaborative. Recommendations were normalized at the drug level and evaluated using recall, case-level AVOID-hit rate, extra-drug burden, precision, Jaccard index, hit@k, and prescription-parameter accuracy.
Results:
Performance differed across workflows. Physician-RAG collaborative showed the highest CORE recall (1.000; 95 % CI, 0.999-1.000), ALT recall (0.955; 95 % CI, 0.945-0.964), overall recall (0.978; 95 % CI, 0.973-0.982), precision (0.963; 95 % CI, 0.957-0.969), and Jaccard index (0.943; 95 % CI, 0.935-0.949). It also had the lowest case-level AVOID-hit rate (0.026; 95 % CI, 0.016-0.039) and the lowest ranked unsafe exposure (AVOID hit@1/3/5: 0.000/0.000/0.008). Prescription-parameter accuracy was highest in the Physician-RAG collaborative workflow (all-correct rate 0.951), compared with 0.948 for the Retrieval-augmented LLM workflow, 0.945 for the Physician workflow, and 0.494 for the Base LLM workflow.
Conclusions:
In this retrospective expert-reference-based evaluation, stand-alone LLM output showed lower agreement and weaker structured prescribing accuracy than supervised workflows. Retrieval augmentation improved coverage and prescribing accuracy, while Physician-RAG collaborative showed the highest agreement with the expert reference standard across completeness, safety screening, prioritization, and structured prescribing quality. These findings support physician-supervised collaboration as a promising direction for medication decision support, but prospective multicenter studies with patient-outcome endpoints are needed before routine clinical deployment.
Related Concept Videos
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Dosage Regimens: Designs and Approaches
Drug Dosing: Geriatric Patients
Drug Dosage Regimen: Overview
Typically, the starting dose and dosing interval are guided by the manufacturer's recommendations based on clinical trials conducted during and after drug...