Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

155
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
155

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The role of endoscope-assisted septectomy and membranectomy for complex chronic subdural hematomas: safety, efficacy, technical feasibility. Patient series.

Journal of neurosurgery. Case lessons·2026
Same author

Supporting Clinical Best Practices after <i>Dobbs</i>.

The New England journal of medicine·2025
Same author

Impact of Donor and Host Age on Systemic Cell Therapy to Treat Age-Related Macular Degeneration.

Cells·2025
Same author

Factors Influencing Postoperative Adherence after Pediatric Kidney Stone Surgery in Alabama (USA): A Single-Institution Retrospective Analysis.

Journal of endourology·2025
Same author

Long-Term Cost Analysis of Initial Panretinal Photocoagulation for Proliferative Diabetic Retinopathy Performed in the Operating Room vs the Clinic.

Journal of vitreoretinal diseases·2025
Same author

Vesicular dermatomyositis presenting without any underlying malignancy.

BMJ case reports·2025

Related Experiment Video

Updated: Jun 8, 2025

Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
05:04

Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training

Published on: August 9, 2024

862

ChatGPT-4 Omni Performance in USMLE Disciplines and Clinical Skills: Comparative Analysis.

Brenton T Bicknell1, Danner Butler2, Sydney Whalen3

  • 1UAB Heersink School of Medicine, 1670 University Blvd, Birmingham, AL, 35233, United States, 1 2566539498.

JMIR Medical Education
|November 6, 2024
PubMed
Summary

The latest ChatGPT-4 Omni model shows significant improvements in medical knowledge and clinical skills, outperforming previous versions and medical students on USMLE-style questions. This highlights its potential as an educational tool for medical training.

Keywords:
AI in medical educationChatGPTChatGPT 3.5ChatGPT 4ChatGPT 4 OmniLLMUSMLEUnited States Medical Licensing Examinationartificial intelligence in medicineclinical skillseducational technologylarge language modelmedical educationmedical licensing examinationmedical student resourcesmedical students

More Related Videos

A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills
07:31

A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills

Published on: February 13, 2020

6.9K
Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
09:52

Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide

Published on: January 15, 2017

17.1K

Related Experiment Videos

Last Updated: Jun 8, 2025

Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
05:04

Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training

Published on: August 9, 2024

862
A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills
07:31

A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills

Published on: February 13, 2020

6.9K
Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
09:52

Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide

Published on: January 15, 2017

17.1K

Area of Science:

  • Artificial Intelligence in Medicine
  • Medical Education Technology
  • Large Language Models

Background:

  • Recent large language models (LLMs) demonstrate capabilities in passing medical licensing exams.
  • A gap exists in detailed analysis of LLM performance across specific medical domains.
  • Assessing LLM utility in medical education requires granular performance data.

Purpose of the Study:

  • To evaluate and compare the accuracy of successive ChatGPT versions (GPT-3.5, GPT-4, GPT-4 Omni).
  • To assess performance across United States Medical Licensing Examination (USMLE) disciplines and clinical clerkships.
  • To evaluate diagnostic and management skills in simulated clinical scenarios.

Main Methods:

  • Utilized 750 clinical vignette-based multiple-choice questions.
  • Tested ChatGPT 3.5, ChatGPT 4, and ChatGPT 4 Omni.
  • Assessed accuracy using a standardized protocol and statistical analysis.

Main Results:

  • GPT-4 Omni achieved 90.4% accuracy, surpassing GPT-4 (81.1%) and GPT-3.5 (60.0%).
  • GPT-4 Omni excelled in social sciences (95.5%), behavioral/neuroscience (94.2%), and pharmacology (93.2%).
  • GPT-4 Omni demonstrated superior diagnostic (92.7%) and management (88.8%) accuracy, significantly outperforming the medical student average (59.3%).

Conclusions:

  • GPT-4 Omni shows substantial improvement in USMLE disciplines and clinical skills compared to predecessors.
  • LLMs like GPT-4 Omni hold significant potential as educational aids for medical students.
  • Careful integration and critical analysis are essential for leveraging LLMs effectively in medical education.