Related Experiment Video
Updated: May 26, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Comparative Analysis of AI-Language Models' MCQ Performance versus Medical Students Across Different Pediatric
Olena Bolgova1, Volodymyr Mavrych1, Eyad Almidani2
1Department of Anatomy, College of Medicine, Alfaisal University, Riyadh, Saudi Arabia.
Advances in Medical Education and Practice
|May 25, 2026
Summary
Large Language Models (LLMs) in pediatric education generally outperform medical students. While advanced AI shows promise, further development is needed for specialized areas like nephrology and neonatology.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
- Pediatric Knowledge Assessment
Background:
- Large Language Models (LLMs) are increasingly adopted in medical education.
- Their efficacy in specialized fields, such as pediatrics, requires further investigation.
Purpose of the Study:
- To evaluate and compare the performance of four leading LLMs on pediatric multiple-choice questions (MCQs).
- To benchmark LLM performance against medical student accuracy and random response baselines.
Main Methods:
- Four LLMs (Copilot, Claude, ChatGPT, Gemini) were tested on 120 pediatric MCQs across six subspecialties.
- LLM consistency was assessed via three attempts per question; medical students and random responses served as benchmarks.
Main Results:
- LLMs achieved an average accuracy of 80.4%, significantly outperforming medical students (72.1%) and random responses (20.5%).
- Copilot (84.5%) and Claude (84.2%) showed the highest accuracy; GPT-4o demonstrated a statistically significant improvement over students.
- LLMs excelled in Pulmonology (96.3%) and Infectious Diseases (85%) but showed lower performance in Nephrology (67.5%) and Neonatology (69.7%).
Conclusions:
- Current LLMs demonstrate strong pediatric knowledge, generally surpassing medical student performance.
- LLMs can be valuable supplementary tools in medical education, but domain-specific improvements are necessary for areas like nephrology and neonatology.