Related Experiment Video
Updated: Feb 11, 2026

Author Spotlight: Simulation and Analysis of the Temperature Rise of Ring Main Unit Equipment
Published on: July 5, 2024
Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational
Dania El Natour1, Mohamad Abou Alfa1, Ahmad Chaaban1
1Faculty of Medicine, Beirut Arab University, Beirut, Lebanon.
New artificial intelligence (AI) models were evaluated on United States Medical Licensing Examination (USMLE) Step 1 questions. Grok achieved the highest accuracy (91.6%), outperforming ChatGPT-4 and other AI tools, especially on image-based questions.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Medical Licensing Examinations
Background:
- Artificial intelligence (AI) models are increasingly integrated into medical education.
- Previous AI models like ChatGPT showed proficiency in USMLE-style questions.
- Emerging AI tools require updated comparative evaluations for accuracy and reliability across medical domains.
Purpose of the Study:
- To compare the performance of five AI models: Grok, ChatGPT-4, Copilot, Gemini, and DeepSeek.
- To assess AI accuracy and consistency on the USMLE Step 1 Free 120-question set.
- To analyze performance across different medical subjects and question formats (text-based, image-based, case-based).
Main Methods:
- Cross-sectional observational study conducted from February 10 to March 5, 2025.
- 119 USMLE-style questions were administered to each AI model using standardized prompts.
- Models answered each question thrice to evaluate consistency; questions were categorized by type; statistical analysis included Chi-square and Fisher's exact tests.
Main Results:
- Grok achieved the highest overall score (91.6%), followed by Copilot (84.9%), Gemini (84.0%), ChatGPT-4 (79.8%), and DeepSeek (72.3%).
- DeepSeek struggled with image-based questions (0% accuracy) but excelled in text-only scenarios (89.6%).
- Grok demonstrated superior performance on image-based (91.3%) and case-based (89.7%) questions, with high consistency (100%).
Conclusions:
- AI models exhibit varied performance across medical domains, with Grok showing top accuracy and consistency.
- Newer models like Grok and Copilot are competitive with established tools like ChatGPT-4.
- Ongoing assessment of AI tools is crucial due to their rapid development and evolving capabilities.
Related Concept Videos
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...
Steps in the Modeling Process
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...
Naturalistic Observations
Rate-Determining Steps
In a multistep reaction mechanism, one of the elementary steps progresses significantly slower than the others. This slowest step is called the rate-limiting step (or rate-determining step). A reaction cannot proceed faster than its slowest step, and hence, the rate-determining step limits the overall reaction rate.
The concept of rate-determining step can be understood from the analogy of a 4-lane freeway with a short-stretch of traffic-bottleneck caused due to...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Observational Learning

