Humans vs. large language models in neurology board examination: performance, limitations, and reference reliability

İlker Arslan1, Müberra Terzi Kumandaş2, Doruk Arslan3

  • 1Department of Neurology, Faculty of Medicine, Hacettepe University, Ankara, Turkey. ilkerarslan94@gmail.com.

Summary

Large language models (LLMs) show high accuracy on neurology board exams, surpassing human examinees. However, their reference reliability is poor, necessitating expert oversight for educational use.

Related Concept Videos