Related Experiment Video
Updated: Aug 10, 2026

A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
USE OF ARTIFICIAL INTELLIGENCE FOR THE EVALUATION OF THE ENTRANCE EXAMINATION OF THE BRAZILIAN SOCIETY OF SHOULDER
Caio Henrique Kenchian1, Karolina Stephany Pereira Ferreira1, Lucas Pereira Sarmento1
1Universidade Federal de Sao Paulo, Escola Paulista de Medicina, Departamento de Ortopedia e Traumatologia, Sao Paulo, SP, Brazil.
Introduction:
AI is increasingly used for medical education and assessment, yet effectiveness on specialized exams remains underexplored. We aim to compare multiple AI models with national human averages on SBCOC entrance exams (2021-2023), evaluate answer accuracy and citation reliability, and contrast models.
Methods:
Five models-ChatGPT-4 (standard and literature-trained), ChatGPT-o1-pro, Gemini, and Meta Llama 3.1-answered official SBCOC exams (50 items/year) individually using a unified prompt and identical images. Each response included one choice (A-D) and a cited source. Outcomes were accuracy and source reliability (peer-reviewed articles/textbooks versus websites). Statistics used chi-square tests and one-way ANOVA with Tukey post-hoc (p<0.05).
Results:
Across 150 questions, ChatGPT-o1-pro led (66%, 62%, 68%) and exceeded human national averages each year (p<0.05). Gemini (40%, 28%, 38%) and Meta Llama 3.1 (32%, 44%, 50%) underperformed relative to humans, while both ChatGPT-4 versions hovered near the ≥50% pass threshold without significant differences. Regarding sources, ChatGPT-o1-pro and Llama predominantly cited articles or books, ChatGPT-4 alternated between literature and websites, and Gemini did not specify references.
Conclusions:
ChatGPT-o1-pro outperformed the national human average and mostly used credible sources; ChatGPT-4 matched humans, while Gemini and Meta Llama 3.1 lagged. Level of evidence IV; Case series.
More Related Videos
05:49Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
03:07Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025