Evaluating the Performance and Fragility of Large Language Models on the Self-Assessment for Neurological Surgeons

Krithik Vishwanath1,2,3, Anton Alyakin1,4, Mrigayu Ghosh5,6

  • 1Department of Neurological Surgery, NYU Langone Health, New York, New York, USA.

Neurosurgery
|December 8, 2025
PubMed
Summary

Large language models (LLMs) show promise for neurosurgery exams but struggle with distractions. Testing revealed that while some LLMs pass neurosurgery board-like questions, their accuracy significantly drops when irrelevant information is present.