Related Experiment Video
Updated: Jun 27, 2025

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
Exploring the Performance of ChatGPT Versions 3.5, 4, and 4 With Vision in the Chilean Medical Licensing Examination:
Marcos Rojas1, Marcelo Rojas2, Valentina Burgess2
1Graduate School of Education, Stanford University, Stanford, CA, United States.
ChatGPT versions 4 and 4V successfully passed the Chilean medical licensing exam (EUNACOM), outperforming version 3.5. Performance varied by medical specialty, indicating a need for tailored AI training in medical education.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Natural Language Processing in Healthcare
Background:
- OpenAI's ChatGPT models (3.5, 4, and 4V) show global potential in medical examinations and education.
- Limited research exists on ChatGPT's efficacy in non-English medical licensing exams, such as Chile's EUNACOM.
- Evaluating ChatGPT's adaptability to diverse linguistic and cultural contexts is crucial for its medical applications.
Purpose of the Study:
- To assess the performance of ChatGPT versions 3.5, 4, and 4V on the EUNACOM, Chile's national medical knowledge exam.
- To compare the accuracy rates of different ChatGPT versions across various medical disciplines within the EUNACOM.
- To identify potential limitations and areas for improvement in AI models for medical licensing assessments.
Main Methods:
- Utilized three official EUNACOM practice drills (540 questions) from the University of Chile.
- Administered three attempts for each ChatGPT version (3.5, 4, and 4V) on the practice drills.
- Systematically categorized and analyzed response accuracy rates for each attempt and ChatGPT version.
Main Results:
- All ChatGPT versions passed the EUNACOM practice drills.
- ChatGPT versions 4 (79.32% accuracy) and 4V (78.83% accuracy) significantly outperformed version 3.5 (57.53% accuracy).
- Versions 4 and 4V showed highest accuracy in surgery, while version 3.5 excelled in psychiatry; performance varied across medical fields.
Conclusions:
- ChatGPT demonstrates capability in passing the EUNACOM, with notable performance differences between versions.
- AI advancements did not significantly improve performance on image-based questions, suggesting a need for specialized training.
- Findings highlight the potential of AI to augment medical education and cognitive skills, necessitating further research into nuanced AI training.
More Related Videos
11:12Driving Simulation in the Clinic: Testing Visual Exploratory Behavior in Daily Life Activities in Patients with Visual Field Defects
Published on: September 18, 2012
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024