Large Language Models Perform at Chance Level in the Diagnosis of Pediatric Pneumonia Using Chest Radiographs.
Justin Gillette1, Michelle Lu2, Thomas F Heston3,4
1Medical Education and Clinical Sciences, Elson S. Floyd College of Medicine, Washington State University, Spokane, USA.
Cureus
|October 22, 2025
Summary
General-purpose large language models (LLMs) show unreliable performance in diagnosing pediatric pneumonia from chest radiographs (CXRs). These AI tools, including ChatGPT, Claude, Gemini, and Grok, are not yet suitable for unsupervised clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Pediatric Radiology
- Medical Diagnostics
Background:
- Pneumonia is a major global child health concern.
- Chest radiographs (CXRs) are crucial for diagnosing pediatric pneumonia.
- Differentiating bacterial from viral pneumonia on CXRs is challenging.
Purpose of the Study:
- To evaluate the diagnostic performance of general-purpose large language models (LLMs) in identifying pediatric pneumonia on CXRs.
- To assess the reliability of LLMs in distinguishing between bacterial pneumonia, viral pneumonia, and normal CXRs.
- To understand the limitations of current LLMs for unsupervised medical image interpretation.
Main Methods:
- Four publicly available LLMs (ChatGPT, Claude, Gemini, Grok) were tested on 44 pediatric CXRs.
- Images were classified as bacterial pneumonia, viral pneumonia, or normal.
- Each LLM interpreted each image twice; accuracy and internal consistency were measured against human expert consensus.
Main Results:
- Average diagnostic accuracy across all LLMs was 31%, equivalent to chance.
- Accuracy was highest for viral pneumonia (54%) and lowest for normal CXRs (18%).
- Internal consistency ranged from 46% to 71%, indicating unreliable performance; concordance with experts did not exceed 49%.
Conclusions:
- General-purpose LLMs are currently unreliable for diagnosing pediatric pneumonia on CXRs.
- Their low accuracy, especially in ruling out disease, and lack of internal consistency pose risks for unsupervised clinical deployment.
- Future AI tools must be purpose-built, trained on diverse data, and integrated with clinical oversight for safe use.
Related Concept Videos
Pneumonia III: Complications and Assessment
777
Pneumonia poses the potential for numerous complications that warrant consideration. These complications include the following:
777
Assessment of Respiration
1.8K
The respiratory system's basic structures and primary functions lay the foundation for nurses' comprehensive respiratory assessments. This assessment includes subjective and objective data to gauge the patient's respiratory health.
Subjective Assessment: Nurses interview the patient to gather information directly during the subjective assessment. It includes questions about the individual's medical history, medications, and symptoms, focusing on past respiratory conditions like...
Subjective Assessment: Nurses interview the patient to gather information directly during the subjective assessment. It includes questions about the individual's medical history, medications, and symptoms, focusing on past respiratory conditions like...
1.8K
Respiratory System Abnormal Finding I: Inspection and Percussion
766
Respiratory system abnormalities are a significant concern in healthcare due to their potential to indicate underlying severe conditions like Chronic Obstructive Pulmonary Disease (COPD), asthma, and pneumonia. These abnormalities can often be detected through physical examination methods like inspection and percussion.
Inspection Findings
During an inspection, several findings may suggest the presence of respiratory distress or disease. Pursed-lip breathing, where exhalation is slowed by...
Inspection Findings
During an inspection, several findings may suggest the presence of respiratory distress or disease. Pursed-lip breathing, where exhalation is slowed by...
766


