Related Experiment Video
Updated: Jan 11, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Artificial Intelligence Clinical Reasoning in Board-Style Clinical Vignettes: A Comparative Study
Lorela Gjunkshi1, Ledio Gjunkshi2, Kenneth A Quezada2
1College of Medicine, Saba University School of Medicine, The Bottom, NLD.
Four artificial intelligence (AI) large language models (LLMs) were tested for diagnostic accuracy using medical licensing exam questions. Claude Sonnet 4 performed best, showing AI
Area of Science:
- Artificial Intelligence in Medicine
- Medical Diagnostics
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly explored for their potential in healthcare applications.
- Evaluating the diagnostic accuracy of AI platforms is crucial for understanding their utility in medical education and practice.
- United States Medical Licensing Examination (USMLE) Step 1 clinical vignettes offer a standardized assessment for diagnostic reasoning.
Purpose of the Study:
- To assess the diagnostic accuracy of four leading large language model (LLM) artificial intelligence (AI) platforms.
- To evaluate the capability of LLMs in generating primary and differential diagnoses from clinical vignettes.
- To compare the performance of different LLM platforms in a simulated diagnostic reasoning task.
Main Methods:
- Ten USMLE Step 1 clinical vignette questions were used, with answer choices removed to create open-ended diagnostic challenges.
- Four LLM platforms (ChatGPT GPT-4o-mini, Meta AI Llama 4, Google Gemini 2.0 Flash, Claude Sonnet 4) were prompted for primary and differential diagnoses.
- Responses were scored using a rubric assessing diagnostic accuracy, with a maximum score of 20 points per model.
Main Results:
- Claude Sonnet 4 achieved perfect accuracy (20/20, 100%), followed by Google Gemini (19/20, 95%), ChatGPT GPT-4o-mini (17/20, 85%), and Meta AI Llama 4 (13/20, 65%).
- All tested LLMs demonstrated clinically relevant diagnostic reasoning.
- Significant variability in diagnostic prioritization and accuracy was observed across the different AI platforms.
Conclusions:
- Current LLMs show considerable potential as supplementary tools for enhancing diagnostic reasoning and medical training.
- The ability of LLMs to generate accurate diagnoses from complex cases supports their value in clinical decision support and education.
- Variability in performance necessitates careful implementation, addressing ethical concerns like bias and privacy before clinical integration.
More Related Videos
08:36The Immersive Cleveland Clinic Virtual Reality Shopping Platform for the Assessment of Instrumental Activities of Daily Living
Published on: July 28, 2022
05:48The Adventures of Fundi Intervention Based on the Cognitive and Emotional Processing in Attention Deficit Hyperactive Disorder Patients
Published on: June 12, 2020
Related Concept Videos
Critical Thinking II
Patient-centered Care
Critical Thinking I
Reason and Intuition
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...