Related Experiment Video
Updated: May 20, 2025

Translational Orthotopic Models of Glioblastoma Multiforme
Published on: February 17, 2023
Accuracy and quality of ChatGPT-4o and Google Gemini performance on image-based neurosurgery board questions
Suyash Sau1, Derek D George2, Rohin Singh2
1University of Rochester School of Medicine and Dentistry, Rochester, NY, USA. suyash_sau@urmc.rochester.edu.
Abstract:
Large-language models (LLMs) have shown the capability to effectively answer medical board examination questions. However, their ability to answer imagebased questions has not been examined. This study sought to evaluate the performance of two LLMs (GPT-4o and Google Gemini) on an image-based question bank designed for neurosurgery board examination preparation. The accuracy of LLMs was tested using 379 image-based questions from The Comprehensive Neurosurgery Board Preparation Book: Illustrated Questions and Answers and Neurosurgery Practice Questions and Answers. LLMs were asked to answer all questions on their own and provide an explanation for their chosen answer. The problem-solving order of questions and quality of LLM responses was evaluated by senior neurological surgery residents who have passed the American Board of Neurological Surgery (ABNS) primary examination. First order questions assess anatomy, second-order questions require diagnostic reasoning, and third-order questions test deeper clinical knowledge by inferring diagnoses and related facts, evaluating the model's ability to recall and apply medical concepts. Chi-squared tests and independent-samples t-tests were conducted to measure performance differences between LLMs. On the image-based question bank, GPT-4o and Gemini achieved correct score percentages of 51.45% (95% CI: 46.43-56.44%) and 39.58% (95% CI: 34.78-44.58%), respectively. GPT-4o significantly outperformed Gemini overall (P = 0.0013), particularly in pathology/histology (P = 0.036) and radiology (P = 0.014). GPT-4o also performed better on second-order questions (56.52% vs. 41.85%, P = 0.0067) and had a higher average response quality rating (2.77 vs. 2.31, P = 0.000002). On a question bank with 379 image-based questions designed for neurosurgery board preparation, GPT-4o obtained a score of 51.45% and outperformed Gemini. GPT-4o not only achieved higher accuracy but also provided higher-quality responses compared to Gemini. In comparison to previous studies on LLM performance of board-style questions, image-based question performance was lower, indicating LLMs may struggle with machine vision/medical image interpretation tasks.
More Related Videos
09:41A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
13:12Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019