Related Experiment Video
Updated: Aug 26, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Navigating uncertainty matters: Evaluating large language models for drug-drug interaction identification
Adeleine Tilley1, Brian Murray1, Kelli Henry2
1Department of Clinical Pharmacy, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of Colorado Anschutz Medical Campus, Aurora.
Background:
Accurate detection of drug-drug interactions (DDIs) is a fundamental component of safe medication management. Traditional rule-based clinical decision support systems for DDI identification lack higher-order reasoning and contribute to alert fatigue. Large language models (LLMs) have potential for DDI identification but may hallucinate, inconsistently identify interactions, and provide overly confident responses despite uncertainty. Prior studies have emphasized accuracy, but few have examined whether LLM uncertainty expression aligns with error risk.
Objective:
To evaluate LLM performance in DDI identification using a clinician-validated dataset and to assess whether prompt-based mitigation strategies improve knowledge-aware uncertainty expression, defined as alignment between refusal behavior and likelihood of error.
Methods:
We developed a clinician-curated DDI identification task consisting of 250 medication lists, each containing 1 clinically relevant interacting drug pair, to evaluate 5 LLMs: GPT-5-Chat, GPT-4o-mini, Gemma-27B, LLaMA3-70B, and Qwen3-32B. Models were evaluated using 3 prompt formats and a zero-shot approach, with no task-specific training or examples provided. Prompts were designed to encourage uncertainty acknowledgment, including a patient safety-focused mitigation prompt to support clinically appropriate and cautious responses. Each case was run 9 times per prompt condition. The primary outcome was the Refusal Index (RI), which quantifies alignment between model refusal behavior and likelihood of error. Secondary outcomes included overall accuracy, accuracy given attempted, refusal rate, self-consistency, F score, weighted score, and entropy.
Results:
Across models, overall DDI identification accuracy ranged from 54.1% to 83.7%. GPT-5-Chat demonstrated the highest overall accuracy and self-consistency, whereas Quen3-32B demonstrated the lowest accuracy but the highest refusal rates. Alignment between refusal behavior and likelihood of error was weak to moderate (RI range 0.104-0.574) and varied by model. Prompt-based mitigation strategies produced inconsistent effects on RI and did not reliably recalibrate uncertainty behavior. Notably, higher overall accuracy and response stability did not consistently correspond to stronger knowledge-aware uncertainty expression. Qwen3-32B and GPT-4o-mini increased refusal rates under mitigation prompting but not in situations where responses were more likely to be incorrect.
Conclusions:
Substantial variability exists in both DDI identification performance and uncertainty calibration across LLMs. Prompt design alone was insufficient to consistently improve knowledge-aware uncertainty expression. Because safe deployment of LLMs in medication management depends not only on accuracy but also on appropriate deferral when error risk is elevated, multidimensional evaluation frameworks are essential before clinical use.
Related Concept Videos
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Quantitative Aspects of Drug-Receptor Interaction
Pharmacogenomics: Identification of New Drug Targets
Drug Discovery: Overview
Pharmacogenetics of Drug Metabolism: Overview
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.