Related Experiment Video
Updated: Aug 6, 2026

Simulator Training for Endovascular Neurosurgery
Published on: May 6, 2020
Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions
Anton Alyakin1,2, Jaden Stryker1, Daniel Alexander Alber1
1Department of Neurological Surgery, NYU Langone Health, New York, New York, USA.
Background And Objectives:
Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions.
Methods:
We developed an automated pipeline using OpenAI generative pre-trained transformer (GPT)-4o and Anthropic Claude Sonnet-3.5 to generate neurosurgical board-style questions from Neurosurgery Publications articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian.
Results:
SANS questions were more often perceived as human-made than GPT-generated (residents, P = .0002; attendings, P = .1091) and Claude-generated (residents, P = .0002; attendings, P = .0272) questions. Notably, 54% of AI-generated questions misled at least one evaluator, and 23% misled both. In quality assessments, SANS questions outperformed GPT-generated (residents, P < 10-5; attendings, P = .0001) and Claude-generated (residents and attendings, P < 10-5) questions. Particularly, 25% of AI-generated questions were rated as suitable for board examinations vs 72% of human-generated questions when measured by evaluator consensus (P < 10-7).
Conclusion:
Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
