Related Experiment Video
Updated: Mar 6, 2026

08:11
Surgical Training for the Implantation of Neocortical Microelectrode Arrays Using a Formaldehyde-fixed Human Cadaver Model
Published on: November 19, 2017
12.0K
A Comparative Study of Large Language Models in Turkish Neurosurgery Education Using a Mock Neurosurgery Board
Kivanc Yangi1, Egemen Gok, Jiuxu Chen
1Barrow Neurological Institute St. Joseph?s Hospital and Medical Center, The Loyal and Edith Davis Neurosurgical Research Laboratory, Arizona, USA.
Turkish Neurosurgery
|March 5, 2026
Summary
Large language models (LLMs) show high accuracy on a mock neurosurgery exam, outperforming residents. Deepseek-R1 and Gemini-2.0 Pro offer valuable educational potential for neurosurgical training.
Area of Science:
- Neurosurgery
- Artificial Intelligence
- Medical Education
Background:
- Large language models (LLMs) are increasingly utilized in medicine.
- The impact of LLMs on neurosurgical education remains understudied.
- This study evaluates the performance of specific LLMs on a neurosurgery board examination.
Purpose of the Study:
- To assess the accuracy and educational value of Deepseek-R1, Gemini-2.0 Pro, ChatGPT-o3-mini-high, and GPT-4.5.
- To compare LLM performance against senior neurosurgery residents.
- To evaluate resident preferences for LLM-generated responses.
Main Methods:
- A 50-question mock neurosurgery board examination was developed.
- The examination was administered to three major LLMs and 10 senior residents.
- Responses were evaluated for accuracy, reasoning time, word count, and readability.
- Residents ranked the educational value of LLM responses.
Main Results:
- All evaluated LLMs (Deepseek-R1, ChatGPT-o3-mini-high, Gemini-2.0 Pro) scored higher in overall accuracy than residents (84%, 82%, 78% vs. 58%, respectively).
- Deepseek-R1 demonstrated the highest accuracy and organized responses, while Gemini-2.0 Pro provided detailed, readable answers.
- Residents preferred Deepseek-R1 and Gemini-2.0 Pro responses over ChatGPT-o3-mini-high.
- GPT-4.5 achieved 74% accuracy, with faster response times and more complex outputs than ChatGPT-o3-mini-high.
Conclusions:
- LLMs demonstrate significant potential as supplementary educational tools in neurosurgical training due to their high accuracy.
- Deepseek-R1 and Gemini-2.0 Pro show promise as refined educational guides or assessment tools for neurosurgery.
- Further development could enhance LLMs for constructing board questions and training assessments.

