Related Experiment Video
Updated: Jun 17, 2026

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Comparing Commercial and Open-Source Large Language Models for Labeling Chest Radiograph Reports
Felix J Dorfner1, Liv Jürgensen1, Leonhard Donle1
1From the Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, 149 Thirteenth St, Charlestown, MA 02129 (F.J.D., T.R.B., M.C.C., A.E.K., C.P.B.); Department of Radiology, Charité-Universitätsmedizin Berlin, corporate member of Freie Universität Berlin and Humboldt Universität zu Berlin, Berlin, Germany (F.J.D., L.D., F.A.M., F.B., L.J.); Department of Pediatric Oncology, Dana-Farber Cancer Institute, Boston, Mass (L.J.); Department of Diagnostic and Interventional Radiology, Technical University of Munich, Munich, Germany (L.C.A.); Mass General Brigham Data Science Office, Boston, Mass (J.S., T.S., C.P.B.); Microsoft Health and Life Sciences (HLS), Redmond, Wash (J.M.); Klinikum rechts der Isar, Technical University of Munich, Munich, Germany (K.K.B.); Department of Radiology and Nuclear Medicine, German Heart Center Munich, Munich, Germany (K.K.B.); and Department of Cardiovascular Radiology and Nuclear Medicine, Technical University of Munich, School of Medicine and Health, German Heart Center, TUM University Hospital, Munich, Germany (K.K.B.).
GPT-4 slightly outperformed open-source large language models (LLMs) in zero-shot chest radiograph report labeling. However, few-shot prompting with examples narrowed the performance gap, showing comparable results for open-source LLMs.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Radiology
- Machine Learning in Healthcare
Background:
- Large language models (LLMs) are rapidly advancing, with numerous commercial and open-source options available.
- Previous studies focused on GPT-4 for radiology report analysis, but real-world comparisons with leading open-source LLMs are lacking.
- Accurate extraction of findings from chest radiograph reports is crucial for clinical decision-making.
Purpose of the Study:
- To compare the performance of leading open-source LLMs against GPT-4 for extracting relevant findings from chest radiograph reports.
- To evaluate the effectiveness of zero-shot and few-shot prompting strategies in this task.
Main Methods:
- Retrospective analysis of two independent datasets of free-text chest radiograph reports (ImaGenome and Massachusetts General Hospital).
- Comparison of commercial models (GPT-3.5 Turbo, GPT-4) with open-source models (Mistral-7B, Mixtral-8×7B, Llama 2-13B, Llama 2-70B, Qwen1.5-72B) and CheXbert/CheXpert-labeler.
- Evaluation using zero-shot and few-shot prompting, with performance measured by F1 scores and compared using the McNemar test.
Main Results:
- On the ImaGenome dataset, Llama 2-70B achieved micro F1 scores of 0.97 (zero-shot) and 0.97 (few-shot), closely matching GPT-4's 0.98.
- On the institutional dataset, an ensemble open-source model achieved micro F1 scores of 0.96 (zero-shot) and 0.97 (few-shot), comparable to GPT-4's 0.98 and 0.97.
- GPT-4 demonstrated superiority in zero-shot labeling, but few-shot prompting significantly improved open-source model performance, yielding comparable results.
Conclusions:
- While GPT-4 excelled in zero-shot report labeling, few-shot prompting with minimal examples enabled open-source LLMs to achieve performance levels close to GPT-4.
- The effectiveness of few-shot prompting varied across different datasets and LLM architectures.
- Open-source LLMs show significant potential for clinical applications in radiology report analysis, especially when fine-tuned with few-shot learning.
More Related Videos
02:09Multi-modal Pulmonary Imaging: Using Complementary Information from CT and Hyperpolarized 129Xe MRI to Evaluate Lung Structure-Function
Published on: April 12, 2024
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Molecular Models
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...