Related Experiment Video
Updated: Jul 18, 2025

12:43
A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
34.8K
Chat GPT as a Neuro-Score Calculator: Analysis of a Large Language Model's Performance on Various Neurological Exam
Tse Chiang Chen1, Emily Kaminski2, Laila Koduri2
1Department of Neurology, Tulane University School of Medicine, New Orleans, Louisiana, USA.
World Neurosurgery
|August 27, 2023
Summary
ChatGPT shows potential in evaluating neurological exams using scales like Glasgow Coma Scale (GCS), Intracranial Hemorrhage (ICH), and Hunt & Hess (H&H). However, its accuracy varies with prompt complexity, indicating limitations for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Natural Language Processing
Background:
- ChatGPT, a large language model, is being explored for medical applications.
- Neurological assessment scales are crucial for patient evaluation.
- This study investigates ChatGPT's performance in applying these scales.
Purpose of the Study:
- To assess the accuracy and reliability of ChatGPT in evaluating patient neurological exams.
- To quantify errors in ChatGPT's application of the Glasgow Coma Scale (GCS), Intracranial Hemorrhage (ICH) score, and Hunt & Hess (H&H) classification.
- To determine the impact of prompt complexity on ChatGPT's performance.
Main Methods:
- Developed 20 patient test cases with detailed neurological exams.
- Created variations of test cases with increasing complexity.
- Assessed ChatGPT's repeatability and quantified errors (Average Error Rate - AER, Magnitude of Errors - AME) for GCS, ICH, and H&H scores.
- Utilized specific prompts for each assessment scale.
Main Results:
- For GCS, the Average Error Rate (AER)/Magnitude of Errors (AME) was 10%/0.150 for base cases.
- Accuracy decreased with complexity; AER reached 45% for complex cases.
- H&H score yielded AER/AME of 13%/0.13, and ICH score yielded 27.5%/0.325.
- Simple prompts resulted in a higher error rate of 70%.
Conclusions:
- ChatGPT demonstrates foundational ability in evaluating neuroexams with established scales (GCS, ICH, H&H).
- Limitations include variable accuracy and potential for 'hallucinations' with complex or vague inputs.
- Further development is needed, but ChatGPT shows promise for future medical applications.

