Related Experiment Video
Updated: Aug 7, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Classifying Usual Interstitial Pneumonia from Radiology Reports: Native Reasoning Versus
Ran Zhang1,2, Thomas M Grist3,4, Mark Schiebler3
1Department of Radiology, School of Medicine and Public Health, University of Wisconsin-Madison, 600 Highland Avenue, Madison, WI, 53792, USA. rzhang229@wisc.edu.
Journal of Imaging Informatics in Medicine
|August 5, 2026
Summary
Large language models (LLMs) show varied performance in classifying usual interstitial pneumonia (UIP) from radiology reports. Prompting strategies must match model architecture for optimal results in clinical AI tasks.
Area of Science:
- Artificial Intelligence in Medicine
- Radiology Informatics
- Natural Language Processing
Background:
- Extracting disease labels from radiology reports is crucial for AI diagnostics and clinical research.
- Classifying usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports is complex, requiring nuanced interpretation.
Purpose of the Study:
- To evaluate the performance of open-source large language models (LLMs) on classifying UIP patterns from HRCT reports.
- To determine if prompting strategies and model architectures influence classification accuracy.
- To assess the impact of model size and native reasoning capabilities on performance.
Main Methods:
- Ten open-source LLMs (Llama, Qwen, Gemma) were tested with three prompting strategies on 270 HRCT reports.
- Expert radiologist consensus was used for ground truth classification.
- Four models were additionally tested in a native reasoning ('thinking') mode.
Main Results:
- The best configuration achieved a Cohen's kappa of 0.70 and 82% four-class accuracy.
- Structured reasoning prompts improved Llama models but degraded models with native reasoning.
- Model size and architecture significantly impacted performance, with larger models not consistently outperforming smaller ones.
Conclusions:
- General LLM improvements do not guarantee enhanced performance on specialized medical tasks like UIP classification.
- Prompt design must be tailored to specific model architectures for optimal clinical application.
- Architecture-dependent interactions between prompting strategies and model capabilities were observed.
