Related Experiment Video
Updated: Jun 27, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Evaluating the Performance and Fragility of Large Language Models on the Self-Assessment for Neurological Surgeons
Krithik Vishwanath1,2,3, Anton Alyakin1,4, Mrigayu Ghosh5,6
1Department of Neurological Surgery, NYU Langone Health, New York, New York, USA.
Large language models (LLMs) show promise for neurosurgery exams but struggle with distractions. Testing revealed that while some LLMs pass neurosurgery board-like questions, their accuracy significantly drops when irrelevant information is present.
Area of Science:
- Artificial Intelligence in Medicine
- Neurosurgery Education
- Natural Language Processing
Background:
- Neurosurgery residents use the Congress of Neurological Surgeons Self-Assessment for Neurological Surgeons for board exam preparation.
- Large language models (LLMs) are increasingly evaluated for their neurosurgical knowledge.
- Clinical text is susceptible to distractions from generative AI and ambient dictation.
Purpose of the Study:
- To assess the performance of state-of-the-art LLMs on neurosurgery board-like questions.
- To evaluate the robustness of LLMs to in-text distractor statements.
Main Methods:
- Evaluated 28 LLMs using 2904 neurosurgery board examination questions.
- Introduced a distraction framework with irrelevant statements containing polysemous words.
- Assessed the degradation of model performance due to distractors.
Main Results:
- Six of 28 LLMs achieved board-passing scores.
- Distractions reduced accuracy by up to 20.4%, causing one model to fail.
- Open-source models showed greater performance decline than proprietary variants.
Conclusions:
- Current LLMs can answer neurosurgery board-like questions but are vulnerable to distractions.
- Developing mitigation strategies is crucial for safe clinical deployment of LLMs.
- LLM resilience against in-text distractions needs enhancement.
More Related Videos
13:12Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Self-Evaluation Maintenance Model