Related Experiment Video
Updated: Jun 15, 2025

11:06
Human Liver Microphysiological System for Assessing Drug-Induced Liver Toxicity In Vitro
Published on: January 31, 2022
4.4K
Toward an Explainable Large Language Model for the Automatic Identification of the Drug-Induced Liver Injury
Chunwei Ma1, Russell D Wolfinger1
1JMP Statistical Discovery, LLC, Cary, North Carolina 27513, United States.
Chemical Research in Toxicology
|August 27, 2024
Summary
Large language models (LLMs) can now automatically identify drug-induced liver injury (DILI) literature. A new LLaMA-2 model achieves high accuracy, improving drug safety analysis and regulatory review.
Area of Science:
- Pharmacovigilance and Computational Toxicology
- Natural Language Processing in Biomedical Research
- Drug Safety and Regulatory Science
Background:
- Drug-induced liver injury (DILI) is a major cause of acute liver failure and a significant challenge in drug safety monitoring.
- Traditional methods like keyword searching and manual curation are insufficient for comprehensive DILI literature retrieval.
- Existing NLP and deep learning models have shown suboptimal performance in identifying DILI-related publications for clinical and regulatory use.
Purpose of the Study:
- To develop and evaluate the first Large Language Model (LLM) specifically designed for DILI literature analysis.
- To leverage the capabilities of LLaMA-2 for automated, high-throughput identification of DILI-related scientific publications.
- To demonstrate the effectiveness of LLMs in enhancing drug safety monitoring and regulatory science.
Main Methods:
- Development of a specialized DILI analysis LLM using the LLaMA-2 architecture.
- Training the model on a large-scale public dataset (14,203 publications) from the CAMDA 2022 literature AI challenge.
- Evaluation using 3-fold cross-validation, comparing performance against smaller models like BERT and GPT variants.
Main Results:
- The LLaMA-2 based LLM achieved an out-of-fold accuracy of 97.19% and an Area Under the ROC Curve (AUC) of 0.9947.
- Demonstrated superior performance compared to other language models in identifying DILI-relevant literature.
- Successfully adapted LLMs, initially for dialogue, into accurate classifiers for biomedical literature analysis.
Conclusions:
- LLMs, particularly LLaMA-2, show significant promise for automated and accurate identification of DILI literature.
- This specialized LLM can facilitate high-throughput analysis, aiding drug safety monitoring and regulatory review processes.
- The study highlights the potential of LLMs to advance regulatory science and improve the efficiency of drug safety evaluations.

