Related Experiment Video
Updated: Apr 28, 2026

An Intestine/Liver Microphysiological System for Drug Pharmacokinetic and Toxicological Assessment
Published on: December 3, 2020
Is ChatGPT Ready for Public Use in Organ-Specific Drug Toxicity Research?
Skylar Connor1, Leihong Wu1, Ruth A Roberts2
1National Center for Toxicological Research, US Food and Drug Administration, Jefferson, AR 72079, USA.
Abstract:
The growing impact of large language models (LLMs), such as ChatGPT, prompts questions about the reliability of their application in public health. We compared drug toxicity assessments by GPT-4 for liver, heart, and kidney against expert assessments using US Food and Drug Administration (FDA) drug-labeling documents. Two approaches were assessed: a 'General prompt', mimicking the conversational style used by the general public, and an 'Expert prompt' engineered to represent an approach of an expert. The Expert prompt achieved higher accuracy (64-75%) compared with the General prompt (48-72%), but the overall performance was moderate, indicating that caution is needed when using GPT-4 for public health. To improve reliability, an advanced framework,such as Retrieval Augmented Generation (RAG), might be required to leverage knowledge embedded in GPT-4.
Related Concept Videos
Preclinical Development: Overview
Toxicity Testing in Animals

