Related Experiment Video
Updated: Jun 30, 2026

A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills
Published on: February 13, 2020
GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination:
Pau Benito1, Mikel Isla-Jover2, Pablo González-Castro3
1Department of Preventive Medicine and Epidemiology, Clinical Institute of Medicine and Dermatology (ICMiD), Hospital Clínic de Barcelona, Rosselló, 138, ground floor, Barcelona, 08036, Spain, 34 932 27 54 00 ext 4046.
Generative artificial intelligence models GPT-4o and OpenAI o1 demonstrate high accuracy on the Spanish medical licensing exam (MIR 2024), outperforming average candidates. These large language models show potential to transform medical education and exam preparation.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Natural Language Processing
Background:
- Generative artificial intelligence (AI) and large language models (LLMs) are rapidly advancing.
- LLMs show significant potential to transform medical education.
- Previous studies have evaluated chatbot performance on medical examinations.
Purpose of the Study:
- To assess the performance of GPT-4o and OpenAI o1 on the Médico Interno Residente (MIR) 2024 examination.
- To compare LLM performance against medical specialists and average candidates on a national medical test.
- To analyze LLM performance across various medical subjects and question difficulties.
Main Methods:
- 176 questions from the MIR 2024 examination were analyzed.
- Each question was presented individually to GPT-4o and OpenAI o1 to ensure independence.
- Accuracy was measured against official answers, with response consistency verified.
- Performance was benchmarked against a consensus of medical specialists and average MIR candidates.
Main Results:
- GPT-4o achieved 89.8% accuracy (90% with verification), and OpenAI o1 achieved 92.6% (93.2% with verification).
- Both LLMs and medical specialists significantly outperformed the average MIR candidate (56.6%).
- LLM accuracy decreased with increasing question difficulty and was slightly higher for clinical cases and positive questions.
Conclusions:
- GPT-4o and OpenAI o1 exhibit excellent and consistent accuracy on the MIR 2024 examination.
- LLMs offer promising opportunities for medical education and preparation for licensing exams.
- Future research should investigate factors influencing LLM accuracy and evaluate emerging AI models.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025