Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large
Matteo Haupt1, Martin H Maurer1
1Department of Diagnostic and Interventional Radiology, Carl von Ossietzky Universität Oldenburg, Oldenburg, Germany (M.H., M.H.M.).
Academic Radiology
|August 12, 2026
Summary
Multimodal large language models (LLMs) show significant performance gaps in pediatric pneumonia detection compared to adults. A domain-trained convolutional neural network (CNN) demonstrated more robust performance across age groups.
Area of Science:
- Radiology and Medical Imaging
- Artificial Intelligence in Healthcare
- Computer Vision
Background:
- Multimodal large language models (LLMs) are emerging tools for image-based radiology tasks.
- Their diagnostic accuracy across diverse patient populations, particularly pediatric versus adult cohorts, is not well understood.
- Evaluating age-related performance differences is crucial for safe clinical integration.
Purpose of the Study:
- To quantify age-related differences in zero-shot LLM performance for pneumonia detection on pediatric versus adult chest radiographs.
- To compare these generalization gaps with a domain-trained convolutional neural network (CNN) baseline.
- To assess the robustness of LLMs and CNNs in distinct age groups.
Main Methods:
- Three leading multimodal LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Pro) were evaluated zero-shot on balanced pediatric and adult chest radiograph test sets.
- Cohort-specific InceptionV3 CNNs were trained and evaluated on the same test sets.
- Performance was measured using the Matthews correlation coefficient (MCC), with domain shift quantified as the difference between adult and pediatric performance (ΔMCC).
Main Results:
- In pediatric cases, the CNN significantly outperformed all LLMs (MCC 0.799 vs. 0.272-0.484).
- While all models improved in adults, the CNN maintained superior performance (MCC 0.850).
- LLMs exhibited larger age-related performance gains (ΔMCC +0.220 to +0.466) compared to the CNN (ΔMCC +0.051), primarily due to improved specificity in adults.
Conclusions:
- Zero-shot multimodal LLMs demonstrate substantial age-related generalization gaps and error asymmetries in pediatric chest radiography.
- Domain-trained CNNs show greater robustness within their specific training domains.
- Thorough subgroup evaluation, especially for pediatric populations, is imperative before the clinical deployment of multimodal LLMs.
