Related Experiment Video
Updated: May 12, 2025

Author Spotlight: Enhancing Transplantation Research Through MicroCT Angiography in Murine Models
Published on: September 22, 2023
Large Language Models for Diagnosing Focal Liver Lesions From CT/MRI Reports: A Comparative Study With Radiologists
Liuji Sheng1, Yidi Chen1,2, Hong Wei1
1Department of Radiology and Functional and Molecular Imaging Key Laboratory of Sichuan Province, West China Hospital, Sichuan University, Chengdu, Sichuan, China.
Two-step ChatGPT-4o demonstrated diagnostic accuracy for focal liver lesions comparable to radiology reports and junior radiologists. However, it was less accurate than experienced radiologists and offered minimal additional benefit when assisting them.
Area of Science:
- Medical Imaging and Diagnostics
- Artificial Intelligence in Healthcare
- Radiology and Hepatology
Background:
- The integration of large language models (LLMs) into the diagnostic workflow for focal liver lesions (FLLs) requires further investigation.
- Assessing the diagnostic capabilities of LLMs like ChatGPT-4o and Gemini against human expertise is crucial for clinical adoption.
Purpose of the Study:
- To evaluate the diagnostic accuracy of two LLMs (ChatGPT-4o and Gemini) in identifying FLLs using CT/MRI reports.
- To compare LLM diagnostic performance with radiologists of varying experience levels and assess their combined utility.
Main Methods:
- A retrospective study analyzed 228 patients with single FLLs diagnosed via CT/MRI and histopathology.
- LLMs were prompted with clinical data and radiology report findings to generate differential diagnoses.
- Radiologists (junior and middle-level) reviewed images independently and with LLM assistance, with performance compared against histopathology.
Main Results:
- Two-step ChatGPT-4o achieved 78.9% accuracy, outperforming single-step LLM use but comparable to radiology reports (80.0%) and junior radiologists (78.9%-82.0%).
- LLM accuracy was lower than middle-level radiologists (84.6%-85.5%).
- ChatGPT-4o did not significantly improve radiologists' diagnostic accuracy when used as an assistant.
Conclusions:
- Two-step ChatGPT-4o shows potential for FLL diagnosis, matching current radiology report accuracy and junior radiologist performance.
- Middle-level radiologists remain superior in diagnostic accuracy for FLLs.
- LLMs currently offer limited incremental value to experienced radiologists in FLL diagnosis.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
06:09Author Spotlight: Advancing Hepatic Fibrosis Diagnosis Using Magnetic Resonance Elastography and AI
Published on: July 21, 2023