Related Experiment Video
Updated: Jan 26, 2026

DUCT: Double Resin Casting followed by Micro-Computed Tomography for 3D Liver Analysis
Published on: September 28, 2021
Characterising liver lesions from free-text computer tomography reports - A real-world multicentre analysis
Jianliang Lu1, Keith Wan-Hang Chiu2, Chelsea Chan3
1Department of Medicine, School of Clinical Medicine, The University of Hong Kong, Hong Kong.
Background:
This study evaluates the performance of a general-purpose (GPT-4) and a medically fine-tuned (Med-LM) large language model (LLM) in classifying liver lesions from unstructured Computed Tomography (CT) reports.
Methods:
Consecutive CT reports (2014-2020) from five institutions were input into GPT-4 and Med-LM withsimple (sp)andoptimised (op) prompts. Lesion-andpatient-level performance were benchmarked against LI-RADS scores assigned by two radiologists, and report quality was analysed using a5-point Likert scale.
Results:
A total of 296 CT reports (mean age, 64.6 years ± 11.3 [SD]; 193 men; 654 lesions) were included. Lesion- and patient-level accuracies for LI-RADS scoring ranged from 40.8% (Med-LMsp) to 61.3% (Med-LMop) and from 27.7% (Med-LMsp) to 52.4% (Med-LMop), respectively. When dichotomized into malignant and benign lesions, lesion- and patient-level accuracies rose to 56.1% (GPT-4sp) - 82.3% (Med-LMop) and 71.3% (Med-LMsp) - 86.5% (Med-LMop). Med-LMop demonstrated the highest performance in all analyses and was statistically superior to other models (all p < 0.001). Non-classification rates ranged between 12.7% (Med-LMop) and 40.5% (GPT-4sp), particularly for benign lesions. Kappa values were weak to moderate between the two reviewers in different aspects of report quality (0.471-0.766), and Likert scores for lesion information differed significantly between correctly and incorrectly classified lesions (all p ≤ 0.04). Repeatability varied widely from 12.7% (Med-LMop) to 39.0% (GPT-4sp).
Conclusions:
Med-LM outperforms GPT-4 in classifying liver lesions from unstructured CT reports with both models better at detecting malignancy than full LI-RADS classification. However, high misclassification rates and inconsistent repeatability hinder their clinical use.
Related Concept Videos
Data Reporting and Recording
Types of Reports I: Hands-off Report
Following are the key components and categories of hand-off reports:
Purpose and Process:
Types of Reports II: Incident or Occurrence Report
Purposes:
In the healthcare industry, reports play a crucial role in documenting incidents within an agency. The primary objective of these reports is to ensure patient safety, uphold the...
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Reporter Genes
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...

