Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparison of a Specialized Large Language Model with GPT-4o for CT and MRI Radiology Report Summarization
Sunyi Zheng1, Nannan Zhao2, Jing Wang3
1Tianjin Medical University Cancer Institute and Hospital, National Clinical Research Center for Cancer, Tianjin's Clinical Research Center for Cancer, State Key Laboratory of Druggability Evaluation and Systematic Translational Medicine, Tianjin Key Laboratory of Digestive Cancer, Key Laboratory of Cancer Prevention and Therapy, Department of Radiology, West Huan-Hu Rd, Ti Yuan Bei, Hexi District, 300060 Tianjin, China.
A specialized large language model (LLM) for radiology report summarization outperformed the general-purpose GPT-4o. LLM-RadSum demonstrated superior performance in generating accurate and clinically useful radiology report summaries.
Area of Science:
- Radiology
- Artificial Intelligence
- Natural Language Processing
Background:
- General-purpose large language models (LLMs) like GPT-4o show potential in radiology.
- The comparative performance of specialized LLMs versus general-purpose LLMs for radiology report summarization is not well-established.
Purpose of the Study:
- To compare the summarization performance of a specialized LLM (LLM-RadSum) against GPT-4o for radiology reports.
- To evaluate the factual consistency, coherence, safety, and clinical utility of generated summaries.
Main Methods:
- Developed LLM-RadSum using a retrospective hospital dataset (training/internal test sets).
- Evaluated F1 scores on internal and external test sets (CT/MRI reports).
- Conducted human evaluation on 1800 reports, comparing LLM-RadSum and GPT-4o outputs using criteria like factual consistency and clinical safety.
Main Results:
- LLM-RadSum achieved higher median F1 scores (0.58) on human evaluation compared to GPT-4o (0.30; P < .001).
- Over 81.5% of LLM-RadSum outputs met expert standards, while 27.8% of GPT-4o outputs required adjustments.
- LLM-RadSum showed superior performance across various modalities, anatomic regions, and patient demographics.
Conclusions:
- A specialized LLM (LLM-RadSum) significantly outperforms general-purpose GPT-4o in summarizing radiology reports.
- LLM-RadSum provides more factually consistent, coherent, safe, and clinically useful summaries.
- Specialized LLMs are more effective for specific medical NLP tasks like radiology report summarization.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
02:09Multi-modal Pulmonary Imaging: Using Complementary Information from CT and Hyperpolarized 129Xe MRI to Evaluate Lung Structure-Function
Published on: April 12, 2024
Related Concept Videos
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Brain Imaging
These technologies include computerized axial tomography (CAT or CT scans), positron-emission tomography (PET scans), magnetic resonance imaging (MRI), functional magnetic resonance imaging (fMRI), and Transcranial Magnetic...
Magnetic Resonance Imaging