Related Experiment Video
Updated: May 23, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Comparative Evaluation of Large Language Models on Pediatric Board-Style Examinations
Nils K T Schönberg1, Daniel J P Deschler2, Julia Hauer1
1Technical University of Munich, Germany; School of Medicine and Health; Department of Pediatrics, Munich, Germany.
Hospital Pediatrics
|May 21, 2026
Summary
Leading large language models (LLMs) perform well on pediatric exams, but struggle with tables. Continued LLM refinement is needed for structured medical data interpretation.
Area of Science:
- Artificial Intelligence in Medical Education
- Natural Language Processing in Healthcare
- Pediatric Board Examinations
Background:
- Large language models (LLMs) are increasingly used in various professional fields.
- Assessing LLM performance on specialized medical examinations is crucial for understanding their capabilities and limitations.
- Pediatric board-style examinations require a high level of medical knowledge and reasoning.
Purpose of the Study:
- To evaluate the performance of four leading LLMs on pediatric board-style examinations.
- To determine if specific question characteristics influence LLM response accuracy.
Main Methods:
- Four LLM platforms were assessed using the 2023 and 2024 American Academy of Pediatrics PREP® Self-Assessment.
- Standardized testing conditions involved sequential input of examination items into each LLM.
- Mixed-effects logistic regression analyzed associations between question characteristics and accuracy.
Main Results:
- All evaluated LLMs exceeded typical human passing thresholds on pediatric board-style exams.
- ChatGPT-5 showed the highest accuracy (89.1-89.6%), followed by Perplexity, Claude Opus 4.1, and Gemini 2.5.
- The presence of tables significantly reduced LLM accuracy (OR 0.52, P=.022), while images and question length did not.
Conclusions:
- Current leading LLMs demonstrate strong, comparable performance on pediatric board exams.
- A shared limitation is difficulty interpreting structured and tabular data, impacting reliability.
- LLMs are valuable for exam preparation but require improvement in handling structured medical information.

