Large Language Models Perform at Chance Level in the Diagnosis of Pediatric Pneumonia Using Chest Radiographs.

Justin Gillette1, Michelle Lu2, Thomas F Heston3,4

  • 1Medical Education and Clinical Sciences, Elson S. Floyd College of Medicine, Washington State University, Spokane, USA.

Cureus
|October 22, 2025
PubMed
Summary

General-purpose large language models (LLMs) show unreliable performance in diagnosing pediatric pneumonia from chest radiographs (CXRs). These AI tools, including ChatGPT, Claude, Gemini, and Grok, are not yet suitable for unsupervised clinical use.