Related Experiment Video
Updated: Apr 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing closed and open large language models on pediatric cardiology board exam performance
Nino Nikolovski1, Conall T Morgan1, Michael N Gritti1
1Division of Cardiology, The Labatt Family Heart Centre, The Hospital for Sick Children, Toronto, Ontario, Canada.
Large language models (LLMs) in pediatric cardiology showed comparable accuracy on board-style questions. These AI tools demonstrate potential as educational aids, though further development is needed for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Pediatric Cardiology Research
Background:
- Large language models (LLMs) are increasingly used in medicine.
- Limited research exists on comparing closed- and open-source LLMs in medical subspecialties.
- Pediatric cardiology requires specialized knowledge, making LLM evaluation in this field crucial.
Purpose of the Study:
- To compare the accuracy of closed-source (ChatGPT-4.0o) and open-source (DeepSeek-R1) large language models.
- To evaluate the performance of these LLMs on a pediatric cardiology board-style examination.
- To discuss the potential educational and clinical utility of LLMs in pediatric cardiology.
Main Methods:
- ChatGPT-4.0o and DeepSeek-R1 were tasked with answering 88 multiple-choice questions from a pediatric cardiology board review textbook.
- Questions covered 11 distinct pediatric cardiology subtopics.
- Processing time for DeepSeek-R1 was recorded and analyzed for correlation with accuracy.
Main Results:
- ChatGPT-4.0o achieved 70% accuracy (62/88), while DeepSeek-R1 achieved 68% accuracy (60/88), with no statistically significant difference (p=0.53).
- Both models performed equally in 5 subtopics, with each outperforming the other in 3 subtopics.
- DeepSeek-R1's processing time showed a negative correlation with accuracy (r=-0.68, p=0.02).
Conclusions:
- ChatGPT-4.0o and DeepSeek-R1 demonstrated comparable accuracy on a pediatric cardiology board examination, nearing the passing threshold.
- The findings suggest that LLMs hold potential as valuable educational tools for pediatric cardiology trainees.
- Further advancements in LLM technology are necessary before widespread clinical integration in pediatric cardiology can be considered.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:55Author Spotlight: Simulating Pediatric Cardiac Surgery Using a Neonatal Piglet Model
Published on: May 26, 2023