Related Experiment Video
Updated: Apr 28, 2026

08:59
An Intestine/Liver Microphysiological System for Drug Pharmacokinetic and Toxicological Assessment
Published on: December 3, 2020
7.8K
Is ChatGPT Ready for Public Use in Organ-Specific Drug Toxicity Research?
Skylar Connor1, Leihong Wu1, Ruth A Roberts2
1National Center for Toxicological Research, US Food and Drug Administration, Jefferson, AR 72079, USA.
Drug Discovery Today
|January 22, 2025
Summary
Large language models (LLMs) show moderate accuracy in assessing drug toxicity for public health applications. Advanced frameworks like Retrieval Augmented Generation may be needed to improve reliability.
Area of Science:
- Artificial Intelligence in Medicine
- Pharmacovigilance
- Computational Toxicology
Background:
- Large language models (LLMs) like ChatGPT are increasingly used, raising concerns about their reliability in public health.
- Accurate drug toxicity assessment is critical for patient safety and regulatory oversight.
Purpose of the Study:
- To evaluate the accuracy of GPT-4 in assessing drug toxicity for liver, heart, and kidney functions.
- To compare the performance of general versus expert prompts for GPT-4 drug toxicity assessments.
Main Methods:
- GPT-4 was prompted using two distinct approaches: a 'General prompt' and an 'Expert prompt'.
- Assessments were benchmarked against expert evaluations derived from US Food and Drug Administration (FDA) drug-labeling documents.
- Drug toxicity for liver, heart, and kidney endpoints was evaluated.
Main Results:
- The 'Expert prompt' demonstrated higher accuracy (64-75%) compared to the 'General prompt' (48-72%).
- Overall performance for GPT-4 in drug toxicity assessment was moderate, suggesting limitations for direct public health application.
- GPT-4's accuracy varied across different organ systems.
Conclusions:
- GPT-4 shows potential but requires cautious application in public health due to moderate accuracy.
- Enhancements, such as Retrieval Augmented Generation (RAG), may be necessary to improve the reliability of LLMs for drug toxicity assessments.
- Further research is needed to optimize LLM frameworks for sensitive public health tasks.
Related Concept Videos
Preclinical Development: Overview
4.8K
Preclinical development consists of a series of tests that ensure the safety and efficacy of a new therapeutic compound before it is tested in humans. There are four main phases to this process. First, safety pharmacology tests are conducted to ensure the drug does not produce any acutely harmful effects. These tests examine parameters such as bronchoconstriction, cardiac dysrhythmias, blood pressure changes, and ataxia. Next, preliminary toxicological testing is performed to determine the...
4.8K
Toxicity Testing in Animals
200
Toxicity tests in animals are grounded on two main assumptions: first, the effects observed in laboratory animals can be extrapolated to humans, especially when adjusted for body surface area; second, high-dose exposure in animals is essential to identify potential human hazards from lower doses. This is based on the quantal dose-response concept, which faces the challenge of extrapolating results from relatively few test animals to much larger human populations. For example, a 0.01% incidence...
200

