Related Experiment Video
Updated: May 10, 2026

09:41
A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
12.3K
Can Artificial Intelligence Mitigate Missed Diagnoses by Generating Differential Diagnoses for Neurosurgeons?
Rohit Prem Kumar1, Vijay Sivan1, Hanin Bachir1
1Department of Neurosurgery, Hackensack Meridian School of Medicine, Nutley, New Jersey, USA.
World Neurosurgery
|May 17, 2024
Summary
Large language models (LLMs) show promise in improving neurosurgical differential diagnoses. While accuracy varies, LLMs can assist in identifying conditions like epilepsy, potentially reducing diagnostic delays.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Neurosurgery
Background:
- Accurate differential diagnoses are critical in neurosurgery.
- Diagnostic delays in neurosurgery lead to significant health and economic challenges.
- Large language models (LLMs) are emerging as potential tools in healthcare.
Purpose of the Study:
- To evaluate the role of LLMs in assisting neurosurgeons with differential diagnoses.
- To assess the diagnostic accuracy of various LLMs in neurosurgical cases.
Main Methods:
- Utilized three chat-based LLMs: ChatGPT (3.5 and 4.0), Perplexity AI, and Bard AI.
- Prompted LLMs with clinical vignettes for 20 neurosurgical disorders.
- Determined LLM accuracy based on correct identification of the target disease within top differentials.
Main Results:
- ChatGPT 3.5 and 4.0 showed initial accuracies of 52.63% and 53.68%, respectively.
- Perplexity AI and Bard AI achieved 40.00% and 29.47% accuracy.
- ChatGPT 3.5 reached 77.89% accuracy for the top 5 differentials; Bard AI improved to 62.11% in the top 5.
- LLMs performed well on common conditions like epilepsy but struggled with complex diseases (e.g., Moyamoya disease).
Conclusions:
- LLMs demonstrate potential to enhance diagnostic accuracy in neurosurgery.
- These AI tools may help decrease the incidence of missed diagnoses.

