Related Experiment Video
Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Comparative Evaluation of Large Language Models on Pediatric Board-Style Examinations
Nils K T Schönberg1, Daniel J P Deschler2, Julia Hauer1
1Technical University of Munich, Germany; School of Medicine and Health; Department of Pediatrics, Munich, Germany.
Objective:
To evaluate the performance of 4 leading large language models (LLMs) on pediatric board-style examinations and to assess whether question characteristics are associated with response accuracy.
Methods:
This study evaluated LLM platforms using the 2023 and 2024 editions of the American Academy of Pediatrics PREP® Self-Assessment. Under standardized testing conditions, all examination items were entered sequentially into each model, and responses were scored for accuracy. Mixed-effects logistic regression was applied to examine associations between question characteristics and correctness, accounting for clustering by question ID.
Results:
Across both examination years, all models achieved high overall accuracy exceeding typical human passing thresholds. ChatGPT-5 demonstrated numerically the highest accuracy (89.1-89.6%), followed by Perplexity (86.1-86.6%), Claude Opus 4.1 (81.8-84.0%), and Gemini 2.5 (78.4-86.3%). Pairwise posthoc comparisons revealed no statistically significant differences among models after adjustment for multiple testing (all adjusted P > .90). Regression analysis identified the presence of tables as a significant predictor of reduced accuracy (odds ratio 0.52, 95% CI 0.29-0.91, P = .022), whereas question length and the presence of images were not significantly associated with performance. No significant interactions between model identity and question characteristics were observed.
Conclusions:
Current market-leading LLMs exhibit strong and broadly comparable performance on pediatric board-style examinations. However, persistent difficulty with structured and tabular data represents a shared limitation that may affect reliability in pediatric educational and clinical contexts. While LLMs are valuable tools for pediatric examination preparation and self-assessment, continued refinement is needed to improve their interpretation of structured medical information.

