Using GPT-4 for LI-RADS feature extraction and categorization with multilingual free-text reports
Kyowon Gu1, Jeong Hyun Lee1, Jaeseung Shin1
1Department of Radiology and Center for Imaging Science, Samsung Medical Center, Sungkyunkwan University School of Medicine, Seoul, Republic of Korea.
Summary
Generative Pre-trained Transformer (GPT)-4 demonstrates high accuracy in extracting Liver Imaging Reporting and Data System (LI-RADS) features from liver MRI reports. Further advancements are needed for reliable clinical application in complex cases.
Area of Science:
- Artificial Intelligence in Radiology
- Medical Imaging Informatics
- Natural Language Processing for Healthcare
Background:
- The Liver Imaging Reporting and Data System (LI-RADS) standardizes hepatocellular carcinoma imaging.
- Radiology report variability hinders automatic data extraction.
- Large language models (LLMs) offer potential for structured data extraction.
Purpose of the Study:
- To evaluate Generative Pre-trained Transformer (GPT)-4 performance in extracting LI-RADS features and categories.
- To assess GPT-4's utility for unstructured free-text liver MRI reports.
Main Methods:
- 160 fictitious Korean/English liver MRI reports were generated by three radiologists.
- Prompt engineering was performed on 20 reports; 140 formed the internal test cohort.
- 72 de-identified real reports constituted the external test cohort; GPT-4 extracted LI-RADS features, and a Python script calculated categories.
Main Results:
- External test accuracy for major LI-RADS features ranged from .92 to .99.
- Accuracy for other LI-RADS features ranged from .86 to .97.
- LI-RADS category extraction accuracy was .85 (95% CI: .76, .93).
Conclusions:
- GPT-4 shows significant promise for extracting LI-RADS features from liver MRI reports.
- Refinement of prompting strategies and neural network architecture is necessary for robust clinical implementation.
- LLMs like GPT-4 could streamline data extraction from complex radiological reports.


