Related Experiment Video
Updated: Apr 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing closed and open large language models on pediatric cardiology board exam performance
Nino Nikolovski1, Conall T Morgan1, Michael N Gritti1
1Division of Cardiology, The Labatt Family Heart Centre, The Hospital for Sick Children, Toronto, Ontario, Canada.
None:
Large language models (LLMs) have gained traction in medicine, but there is limited research comparing closed- and open-source models in subspecialty contexts. This study evaluated ChatGPT-4.0o and DeepSeek-R1 on a pediatric cardiology board-style examination to quantify their accuracy and discuss educational and clinical utility. ChatGPT-4.0o and DeepSeek-R1 were used to answer 88 text-based multiple choice questions across 11 pediatric cardiology subtopics from a Pediatric Cardiology Board Review textbook. DeepSeek-R1's processing time per question was measured. ChatGPT-4.0o and DeepSeek-R1 achieved 70% (62/88) and 68% (60/88) accuracy, respectively (p = 0.53). Subtopic accuracy was equal in 5 of 11 chapters, with each model outperforming its counterpart in 3 of 11. DeepSeek-R1's processing time negatively correlated with accuracy (r = -0.68, p = 0.02). ChatGPT-4.0o and DeepSeek-R1 were comparable in accuracy and approached the passing threshold on a pediatric cardiology board examination. While further development of LLMs is required for clinical integration into pediatric cardiology, these findings suggest the potential utility of these models as educational aids.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:55Author Spotlight: Simulating Pediatric Cardiac Surgery Using a Neonatal Piglet Model
Published on: May 26, 2023