Related Experiment Video
Updated: Jun 5, 2026

13:12
Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Humans vs. large language models in neurology board examination: performance, limitations, and reference reliability
İlker Arslan1, Müberra Terzi Kumandaş2, Doruk Arslan3
1Department of Neurology, Faculty of Medicine, Hacettepe University, Ankara, Turkey. ilkerarslan94@gmail.com.
Acta Neurologica Belgica
|June 4, 2026
Summary
Large language models (LLMs) show high accuracy on neurology board exams, surpassing human examinees. However, their reference reliability is poor, necessitating expert oversight for educational use.
Area of Science:
- Neurology
- Artificial Intelligence
- Medical Education
Background:
- Large language models (LLMs) are increasingly used in various fields, including medicine.
- Evaluating their performance in specialized medical domains like neurology is crucial.
Purpose of the Study:
- To assess the accuracy and reference reliability of three leading LLMs (ChatGPT Plus, Gemini Advanced, Microsoft Copilot Pro) on neurology board examination questions.
- To compare LLM performance against human examinee performance.
Main Methods:
- 803 multiple-choice questions from Turkish National Neurology Board Examinations (2015-2024) were used.
- LLMs were prompted to provide answers and supporting references.
- Performance was analyzed by neurological subspecialty, question type, and visual content, and compared to human examinee data.
- References for incorrect answers were evaluated by neurologists.
Main Results:
- LLMs achieved high accuracy rates (85-87%), significantly outperforming the human average of 65%.
- Accuracy was comparable across LLMs and did not vary by question type or subspecialty.
- LLMs performed better on non-visual questions; no advantage was seen on visual items.
- Significant issues with reference reliability were found, including insufficient citations and fabricated references.
Conclusions:
- LLMs demonstrate strong performance on neurology board exams, exceeding human examinee accuracy.
- Visually based questions and reference quality present limitations for LLMs.
- LLMs should be utilized as supplementary educational tools under expert supervision, not as independent authorities.
Keywords:
Artificial Intelligence (AI)Board examinationLarge language models (LLM)Neurology educationReference reliabilityMore Related Videos
Related Concept Videos
Higher Mental Functions of the Brain: Language
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language and Cognition
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.

