Related Experiment Video
Updated: Feb 8, 2026

Simulator Training for Endovascular Neurosurgery
Published on: May 6, 2020
Comparative diagnostic capability of large language models in neurosurgery
Akshay Warrier1, Yaxel Levin-Carrion1, Shrey B Shah1
11Department of Neurosurgery, Rutgers New Jersey Medical School, Newark.
Objective:
OpenAI, Google, and Microsoft have recently developed popular large language models (LLMs) with incredible clinical applications. LLMs specific to neurosurgery, such as AtlasGPT, have also been recently released. However, the comparative neurosurgical diagnostic capabilities of these models are not well studied. The aim of this study was to evaluate and compare the ability of LLMs to diagnose neurosurgical pathologies.
Methods:
Clinical vignettes (n = 148) extracted from a common neurosurgery case-based review textbook were stratified by subspecialty. OpenAI's ChatGPT-3.5 and ChatGPT-4, Google's Gemini, Microsoft Copilot, and AtlasGPT were prompted to provide a diagnosis: "Provide a neurosurgical diagnosis given the following history…[vignette]." Imaging was inputted for capable LLMs, and all queries were run in May 2024. Diagnoses were compared with the textbook for accuracy and errors were categorized appropriately.
Results:
ChatGPT-4 was the most accurate model (74% correct), followed by AtlasGPT (63% correct), ChatGPT-3.5 (53% correct), Microsoft Copilot (48% correct), and Gemini (36% correct). Chi-square comparisons demonstrated that ChatGPT-4 was more accurate in providing clinical diagnoses than its counterparts (p = 0.005). Across all vignettes and LLMs, most errors were due to an inability to attribute a key piece of information (generally imaging data) to the diagnostic process while otherwise using logical stepwise reasoning.
Conclusions:
ChatGPT-4 offered the most accurate diagnoses when given established clinical vignettes. Adding imaging processing capabilities and relevant data significantly increased the accuracy of LLM diagnoses. LLMs can offer accurate assessments of common neurosurgical conditions but necessitate detailed prompting from clinicians. Artificial intelligence has incredible clinical potential; however, practitioners must be cautious and think critically while using them for diagnostic purposes.
More Related Videos
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...

