Related Experiment Video
Updated: Aug 7, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Classifying Usual Interstitial Pneumonia from Radiology Reports: Native Reasoning Versus
Ran Zhang1,2, Thomas M Grist3,4, Mark Schiebler3
1Department of Radiology, School of Medicine and Public Health, University of Wisconsin-Madison, 600 Highland Avenue, Madison, WI, 53792, USA. rzhang229@wisc.edu.
Journal of Imaging Informatics in Medicine
|August 5, 2026
Summary
Large language models (LLMs) show promise for classifying usual interstitial pneumonia (UIP) from radiology reports. However, performance varies by model architecture and prompting strategy, not just size.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Natural Language Processing in Healthcare
Background:
- Extracting disease labels from radiology reports is crucial for AI diagnostic models and clinical research.
- Classifying usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports is complex, requiring detailed analysis of imaging features and exclusion criteria.
Purpose of the Study:
- To evaluate the performance of open-source large language models (LLMs) on classifying UIP patterns from HRCT reports.
- To investigate how different prompting strategies and model architectures (Llama, Qwen, Gemma) affect classification accuracy.
- To determine if advanced reasoning capabilities or larger model sizes correlate with improved performance on this specialized medical task.
Main Methods:
- Ten open-source LLMs (8-405B parameters) from three families were tested on 270 expert-classified HRCT reports.
- Three prompting strategies were applied, with four models additionally tested in "thinking" mode.
- Performance was assessed using Cohen's kappa and four-class accuracy.
Main Results:
- The best configuration achieved a Cohen's kappa of 0.70 and 82% four-class accuracy.
- Structured reasoning prompts improved Llama models but degraded models with native reasoning (Qwen, Gemma).
- Larger model size did not consistently predict better performance; some smaller models outperformed larger ones.
Conclusions:
- General LLM improvements and larger sizes do not guarantee better performance on specialized medical classification tasks like UIP detection.
- Prompt design must be carefully matched to the specific model architecture for optimal results.
- Architecture-dependent interactions between prompting strategies and model capabilities are critical for clinical NLP applications.
