Related Experiment Video
Updated: Sep 19, 2026

Human Liver Microphysiological System for Assessing Drug-Induced Liver Toxicity In Vitro
Published on: January 31, 2022
Multicenter Evaluation of Large Language Models Versus Hepatologists for Prognostic Prediction in Drug-Induced Liver
Haoshuang Fu1, Yanan Du1, Gangde Zhao2
1Department of Infectious Diseases, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
Background & Aims:
Drug-induced liver injury (DILI) would progress to chronicity or death. Large language models (LLMs) may enhance clinical decision-making, yet their utility relative to physicians in DILI remains unclear. Therefore, we evaluated their performance in predicting DILI outcomes.
Methods:
We enrolled 943 DILI patients from three centers as internal and external cohorts. Based on 12-month follow-up, outcomes were classified as recovery, 6/12-month chronicity, and death. LLMs (Gemini-2.5 Pro, GPT-5.1, DeepSeek-3.2), hepatologists (Junior, middle, senior), and models (Hy's Law, nHy's Law, MELD Score) estimated probabilities of outcomes. LLM-Rules (VOTE, OR, AND) were applied to enhance stability. Model performance was assessed.
Results:
For 6-month chronicity, the senior achieved highest AUROC (0.61) with an accuracy of 70%. Gemini-2.5 Pro and GPT-5.1 yielded AUROCs of 0.60 and 0.59, respectively, outperforming junior and middle hepatologists. Gemini-2.5 Pro demonstrated strongest agreement with senior (κ = 0.43). LLMs all exhibited lower accuracy and specificity than hepatologists. A similar result was observed in 12-month chronicity. For overall mortality, the senior achieved highest AUROC (0.87) with an accuracy of 83%. Gemini-2.5 Pro and GPT-5.1 achieved AUROCs of 0.86, outperforming junior hepatologist, Hy's Law, and nHy's Law. GPT-5.1 achieved strongest agreement with senior (κ = 0.25). LLM-Rules demonstrated stability for predicting outcomes across cohorts. OR and AND rules improved sensitivity and specificity, respectively.
Conclusions:
GPT-5.1 and Gemini-2.5 Pro showed AUROCs approaching senior hepatologists for DILI outcomes with limited accuracy and specificity. LLM-Rules demonstrated stable performance across cohorts with improved sensitivity or specificity, supporting the potential of multi-LLM approaches as clinician-supervised complementary tools.
