Related Experiment Video
Updated: Jan 26, 2026

DUCT: Double Resin Casting followed by Micro-Computed Tomography for 3D Liver Analysis
Published on: September 28, 2021
Characterising liver lesions from free-text computer tomography reports - A real-world multicentre analysis
Jianliang Lu1, Keith Wan-Hang Chiu2, Chelsea Chan3
1Department of Medicine, School of Clinical Medicine, The University of Hong Kong, Hong Kong.
Medically fine-tuned large language models (LLMs) like Med-LM show promise in classifying liver lesions from CT reports, outperforming general models. However, high misclassification rates and inconsistent repeatability currently limit their clinical application.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Clinical Data
- Machine Learning for Diagnostic Support
Background:
- Evaluating large language models (LLMs) for medical applications is crucial.
- General-purpose (GPT-4) and medically fine-tuned (Med-LM) LLMs were assessed.
- Classification of liver lesions from unstructured Computed Tomography (CT) reports was the focus.
Purpose of the Study:
- To compare the performance of GPT-4 and Med-LM in classifying liver lesions.
- To benchmark LLM performance against radiologist-assigned LI-RADS scores.
- To analyze the impact of prompt optimization on LLM classification accuracy.
Main Methods:
- Utilized 296 consecutive CT reports (2014-2020) from five institutions.
- Employed simple (sp) and optimized (op) prompts for GPT-4 and Med-LM.
- Benchmarked lesion- and patient-level performance against radiologist LI-RADS scores and report quality metrics.
Main Results:
- Med-LM, particularly with optimized prompts (Med-LMop), demonstrated superior performance in all analyses (p < 0.001).
- Dichotomizing lesions into malignant and benign categories improved accuracies significantly for both models.
- High non-classification rates (12.7%-40.5%) and variable repeatability (12.7%-39.0%) were observed, especially for benign lesions.
Conclusions:
- Med-LM significantly outperforms GPT-4 in classifying liver lesions from CT reports.
- Both LLMs excel at detecting malignancy compared to full LI-RADS classification.
- Clinical utility is currently limited by high misclassification rates and inconsistent repeatability.
Related Concept Videos
Data Reporting and Recording
Types of Reports I: Hands-off Report
Following are the key components and categories of hand-off reports:
Purpose and Process:
Types of Reports II: Incident or Occurrence Report
Purposes:
In the healthcare industry, reports play a crucial role in documenting incidents within an agency. The primary objective of these reports is to ensure patient safety, uphold the...
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Reporter Genes
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...

