Related Experiment Video
Updated: Jan 8, 2026

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
Comparative evaluation of artificial intelligence platforms and drug interaction screening databases using real-world
Bálint Márk Domián1, Amir Reza Ashraf1, András Tamás Fittler1
1Department of Pharmaceutics, Faculty of Pharmacy, University of Pécs, Pécs, Hungary.
Large language models (LLMs) show potential for drug-drug interaction (DDI) screening, but current AI tools lack the precision and sensitivity for reliable clinical use. Professional validation of LLM outputs is crucial for patient safety in polypharmacy management.
Area of Science:
- Pharmacology
- Artificial Intelligence
- Clinical Pharmacy
Background:
- Polypharmacy increases the risk of harmful drug-drug interactions (DDIs).
- Conventional DDI screening tools often have limited coverage and generate excessive alerts, leading to alert fatigue.
- Large language models (LLMs) present a novel approach for DDI identification, but their real-world clinical utility requires further investigation.
Purpose of the Study:
- To compare the performance of conventional DDI databases with LLM-based screening using real-world patient data.
- To evaluate the sensitivity, specificity, precision, and F1 scores of ChatGPT, Google Gemini, and Microsoft Copilot in identifying clinically relevant DDIs.
Main Methods:
- An exploratory study utilizing anonymized medication lists from rheumatology patients.
- A reference set of 204 clinically relevant DDIs was established using Lexicomp, Medscape, and Drugs.com.
- ChatGPT, Google Gemini, and Microsoft Copilot were queried using identical prompts to identify potential DDIs.
Main Results:
- LLMs identified a high volume of potential DDIs, with Gemini identifying 1556 and Copilot 1813 interactions.
- ChatGPT achieved the highest specificity (0.868), while Gemini had the highest sensitivity (0.697).
- All evaluated LLMs demonstrated low precision, and none achieved the necessary balance of sensitivity and specificity for reliable clinical decision-making, with ChatGPT showing the highest F1 score (0.2520).
Conclusions:
- LLMs show promise in identifying true DDIs but are limited by "hallucinations" producing clinically inaccurate information.
- Current LLMs are not reliable as standalone DDI screening tools and require professional validation.
- LLMs can potentially support clinical pharmacists in managing polypharmacy, but human oversight is essential for patient safety.
More Related Videos
Related Concept Videos
Bioequivalence of Drugs: Drugs with Multiple Indications
Quantitative Aspects of Drug-Receptor Interaction
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Drug Products: Biologics, Biosimilars and Interchangeables

