Related Experiment Video
Updated: May 28, 2025

A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
Performance evaluation of ChatGPT-4.0 and Gemini on image-based neurosurgery board practice questions: A comparative
Alana M McNulty1, Harshitha Valluri1, Avi A Gajjar1
1Department of Neurosurgery, Albany Medical Center, Albany, NY, USA.
Introduction:
Artificial intelligence (AI) has gained significant attention in medicine, particularly in neurosurgery, where its potential is often discussed and occasionally feared. Large language models (LLMs), such as ChatGPT-4.0 (OpenAI) and Gemini (formerly known as Bard, Google DeepMind), have shown promise in text-based tasks but remain under explored in image-based domains, which are essential for neurosurgery. This study evaluates the performance of ChatGPT-4.0 and Gemini on image-based neurosurgery board practice questions, focusing on their ability to interpret visual data, a critical aspect of neurosurgical decision-making.
Methods:
A total of 250 image-based questions selected from two neurosurgical board review books were obtained. Each question was presented to both ChatGPT-4.0 and Gemini in its original format, including images such as MRI scans, pathology slides, and surgical visuals. The models were tasked with answering the questions, and their accuracy was determined based on the number of correct responses.
Results:
ChatGPT-4.0 accurately answered 135/250 (54.0 %) of questions, while Gemini correctly answered 24/250 (8.6 %). ChatGPT accurately answered 85/129 (65.9 %) of The Comprehensive Neurosurgery Board Preparation Book, and 50/121 (41.3 %) of the Neurosurgery Board Review book. Gemini answered 23/129 (17.8 %) of The Comprehensive Neurosurgery Board Preparation Book, and only 1 of 121 (0.8 %) questions from Neurosurgery Board Review. When comparing the ability of ChatGPT-4.0o vs. Gemini, ChatGPT significantly outperformed Gemini in accuracy (p < 0.0001). The overall refusal rate for Gemini in answering questions was 102/250 (40.8 %), while ChatGPT attempted to answer all questions.
Conclusions:
While ChatGPT-4.0 demonstrated some capacity to interpret image-based neurosurgery board questions, both models exhibited significant limitations, particularly in processing and analyzing complex visual data. These findings emphasize the need for targeted advancements in AI to improve visual interpretation in neurosurgical education and practice.

