Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative performance of large language models on cardiovascular certification simulation exam
Eesha Nachnani1, Kashish Goel2, Alexander E Sullivan2
1University School of Nashville, Nashville, TN.
Abstract:
Artificial intelligence (AI) is becoming increasingly prevalent in medical practice and has demonstrated sufficient clinical acumen to pass several licensing examinations. We tested the ability of 3 popular large-language models, ChatGPT-4.0 (OpenAI), Gemini (Google), and Bing AI (Microsoft), to pass a cardiovascular medicine board-style exam. Of these AI platforms, only ChatGPT-4.0 was able to achieve a score similar to human participants.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
