Related Experiment Video
Updated: May 10, 2026

Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
Performance of Publicly Available Large Language Models on Internal Medicine Board-style Questions
Constantine Tarabanis1, Sohail Zahid1, Marios Mamalis2
1Leon H. Charney Division of Cardiology, NYU Langone Health, New York University School of Medicine, New York, New York, United States of America.
Large language models (LLMs) show promise in medical exams, with GPT-4.0 outperforming human physicians on internal medicine board questions. Augmenting LLMs with medical texts improves their accuracy.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Large Language Model (LLM) Evaluation
Background:
- Ongoing research benchmarks large language models (LLMs) against physician knowledge using medical examinations.
- Limited studies have assessed LLM performance on internal medicine (IM) board examination questions.
- The impact of domain-specific knowledge augmentation on LLM performance in medical contexts is not well-established.
Purpose of the Study:
- To evaluate the performance of various LLMs (GPT-3.5, GPT-4.0, LaMDA, Llama 2) on internal medicine board-style questions.
- To assess the effect of input augmentation using medical texts (Harrison's Principles of Internal Medicine) on LLM performance.
- To compare LLM performance against human respondents on IM board examination questions.
Main Methods:
- 240 randomly selected IM board-style questions from the Medical Knowledge Self-Assessment Program (MKSAP) were used.
- LLMs were accessed via API and chatbot interfaces, with and without Retrieval Augmented Generation (RAG) using Harrison's Principles of Internal Medicine.
- LLM-generated explanations for correct answers were compared to human-generated explanations by a blinded, board-certified physician.
Main Results:
- GPT-4.0 achieved the highest scores (77.5-80.7%), outperforming GPT-3.5, human respondents, LaMDA, and Llama 2.
- GPT-4.0 surpassed human MKSAP users across all IM subjects, with notable strengths in Infectious Disease and Rheumatology.
- Input augmentation via RAG increased GPT-3.5 and GPT-4.0 performance by 4.5-7.5% when accessed via API.
Conclusions:
- GPT-4.0 demonstrates superior performance on internal medicine board-style questions compared to other LLMs and human respondents.
- Retrieval Augmented Generation (RAG) is a viable technique for enhancing LLM accuracy in medical examinations by incorporating domain-specific knowledge.
- LLM performance can vary based on access method (API vs. chatbot), with chatbots generally showing higher scores.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024