Related Experiment Video
Updated: May 16, 2026

Whole-brain Segmentation and Change-point Analysis of Anatomical Brain MRI—Application in Premanifest Huntington's Disease
Published on: June 9, 2018
Towards Community-Based Evaluation of AI in Neurology: Development of a Headache Diagnosis Dataset for Large Language
Anika Zahn1, Sebastian Strauss1, Dorian Zwanzig2
1Department of Neurology, University Medicine Greifswald, Greifswald, Germany.
Abstract:
Diagnosing headache disorders remains a clinical challenge due to the heterogeneity of headache phenotypes and the absence of objective biomarkers. This study presents a curated dataset of 50 clinical headache case examples, comprising both real (n = 34) and synthetic (n = 16) cases, categorized across 20 diagnoses according to ICHD-3 criteria. The dataset enables the evaluation of large language models (LLMs) for diagnostic accuracy in headache medicine. Three GPT-based models were tested using different prompting strategies, with diagnostic performance assessed at both diagnosis and group levels. Top-1 accuracy ranged from 24% to 63% at the diagnosis level and up to 92% at the group level. The results highlight the potential of LLMs in supporting differential diagnosis of headache disorders, while also emphasizing the need for further validation with larger, diverse datasets. Future efforts will focus on expanding real-world data through clinical collaborations and benchmarking LLMs against medical professionals to assess their utility in clinical decision-making.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Ethics in Research
The Stanford Prison Experiment
Bullying
Sex Linked Disorders
Sexually Transmitted Infections
Conduct Disorder