Related Experiment Video
Updated: Jun 18, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Stroke Diagnosis and Prediction Tool Using ChatGLM: Development and Validation Study
Xiaowei Song1, Jiayi Wang2, Feifei He3
1Department of Neurology, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua University, Beijing, China.
Insights
A new large language model (LLM) accurately diagnoses stroke using clinical notes and noncontrast computed tomography (NCCT) reports. This AI tool shows high accuracy in identifying stroke types and guiding treatment, potentially reducing patient disability and mortality.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Neurology
Background:
- Stroke is a leading cause of death and disability worldwide.
- Accurate and timely diagnosis is crucial for effective stroke treatment and improved patient outcomes.
- Current stroke diagnosis faces challenges, leading to discrepancies in care.
Purpose of the Study:
- To develop and validate a large language model (LLM) for stroke diagnosis and prediction.
- To integrate free-text electronic health records and noncontrast computed tomography (NCCT) reports for enhanced stroke detection.
- To improve the accuracy and speed of stroke identification and guide recanalization therapy.
Main Methods:
- Utilized the ChatGLM-6B large language model (LLM) for stroke diagnosis.
- Employed instruction tuning and low-rank adaptation (LoRA) techniques for model optimization.
- Trained and validated the LLM on a dataset of 1885 patients and externally tested on 335 patients from multiple hospitals.
Main Results:
- The LLM achieved 99% accuracy in internal validation and up to 95.5% in external validation for stroke diagnosis.
- Demonstrated high accuracy in differentiating ischemic stroke from hemorrhage (up to 100%) and identifying large vessel occlusions (up to 88.6%).
- Showcased effectiveness in screening patients for intravenous thrombolysis (IVT) with accuracies up to 89.4%.
Conclusions:
- An LLM integrating clinical text and NCCT reports can effectively identify strokes and inform recanalization therapy decisions.
- The developed LLM shows significant potential to enhance stroke identification accuracy and reduce critical reperfusion times.
- Further validation through widespread deployment is recommended to confirm the clinical utility of this AI-driven diagnostic tool.
Background:
Stroke is a globally prevalent disease that imposes a significant burden on health care systems and national economies. Accurate and rapid stroke diagnosis can substantially increase reperfusion rates, mitigate disability, and reduce mortality. However, there are considerable discrepancies in the diagnosis and treatment of acute stroke.
Objective:
The aim of this study is to develop and validate a stroke diagnosis and prediction tool using ChatGLM-6B, which uses free-text information from electronic health records in conjunction with noncontrast computed tomography (NCCT) reports to enhance stroke detection and treatment.
Methods:
A large language model (LLM) using ChatGLM-6B was proposed to facilitate stroke diagnosis by identifying optimal input combinations, using external tools, and applying instruction tuning and low-rank adaptation (LoRA) techniques. A dataset containing details of 1885 patients with and those without stroke from 2016 to 2024 was used for training and internal validation; another 335 patients from two hospitals were used as an external test set, including 230 patients from the training hospital but admitted at different periods, and 105 patients from another hospital.
Results:
The LLM, which is based on clinical notes and NCCT, demonstrates exceptionally high accuracy in stroke diagnosis, achieving 99% in the internal validation dataset and 95.5% and 79.1% in two external test cohorts. It effectively distinguishes between ischemia and hemorrhage, with an accuracy of 100% in the validation dataset and 99.1% and 97.1% in the other test cohorts. In addition, it identifies large vessel occlusions (LVO) with an accuracy of 80% in the validation dataset and 88.6% and 83.3% in the other test cohorts. Furthermore, it screens patients eligible for intravenous thrombolysis (IVT) with an accuracy of 89.4% in the validation dataset and 60% and 80% in the other test cohorts.
Conclusions:
We developed an LLM that leverages clinical text and NCCT to identify strokes and guide recanalization therapy. While our results necessitate validation through widespread deployment, they hold the potential to enhance stroke identification and reduce reperfusion time.
More Related Videos
Related Concept Videos
Cancer Survival Analysis
Statistical Software for Data Analysis and Clinical Trials

