Can ChatGPT 4.0 Diagnose Epilepsy? A Study on Artificial Intelligence's Diagnostic Capabilities
Francesco Brigo1, Serena Broggi2, Eleonora Leuci3
1Innovation, Research and Teaching Service (SABES-ASDAA), Teaching Hospital of the Paracelsus Medical Private University (PMU), 39100 Bolzano, Italy.
Artificial intelligence (AI), specifically large language models (LLMs), shows very low agreement with human experts in diagnosing epilepsy. While better at ruling out epilepsy, AI tools require significant improvement for clinical use in epilepsy diagnosis.
Area of Science:
- Neurology
- Medical Informatics
- Artificial Intelligence
Background:
- Epilepsy diagnosis relies on expert interpretation of complex clinical data.
- Artificial intelligence (AI), particularly large language models (LLMs), offers potential for enhancing diagnostic decision support.
- Evaluating AI's diagnostic capabilities against human expertise is crucial for its safe implementation.
Purpose of the Study:
- To compare the diagnostic agreement of ChatGPT 4.0 with human epileptologists for epilepsy diagnosis.
- To assess the performance metrics (sensitivity, specificity) of AI in epilepsy diagnosis.
- To identify factors contributing to diagnostic errors made by ChatGPT.
Main Methods:
- Retrospective analysis of 597 patient records from emergency department visits for seizures.
- Comparison of diagnoses made by epileptologists and ChatGPT 4.0 using 2014 International League Against Epilepsy (ILAE) criteria.
- Statistical analysis including Cohen's kappa, 2x2 contingency tables, and multivariate analysis.
Main Results:
- Epilepsy was diagnosed in 36.2% of patients by neurologists versus 18.2% by ChatGPT.
- Very low agreement was observed between human and AI diagnoses (Cohen's kappa = -0.01).
- ChatGPT demonstrated low sensitivity (17.6%) but higher specificity (81.4%) for epilepsy diagnosis, with errors more common in older patients.
Conclusions:
- ChatGPT 4.0 exhibits significantly lower performance than human clinicians in diagnosing epilepsy.
- AI shows a higher ability to identify non-epileptic cases than epilepsy cases.
- Substantial further research is required to enhance the diagnostic accuracy and clinical utility of LLMs in epilepsy management.
More Related Videos
10:23Equipment Setup and Artifact Removal for Simultaneous Electroencephalogram and Functional Magnetic Resonance Imaging for Clinical Review in Epilepsy
Published on: June 23, 2023
09:41A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
