Related Experiment Video
Updated: Jun 13, 2026

10:28
Dynamic Digital Biomarkers of Motor and Cognitive Function in Parkinson's Disease
Published on: July 24, 2019
15.2K
Early identification of stroke through deep learning with multi-modal human speech and movement data
Zijun Ou1, Haitao Wang1, Bin Zhang2
1School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, Guangdong Province, China.
Neural Regeneration Research
|May 20, 2024
Summary
A new multimodal deep learning model, based on the Face Arm Speech Test (FAST), shows higher clinical value for early stroke identification. This AI tool effectively uses video and audio data to improve stroke assessment accuracy in emergency settings.
Area of Science:
- Neurology
- Artificial Intelligence
- Medical Imaging
Background:
- Early stroke identification is crucial for improving patient outcomes and quality of life.
- Current clinical stroke screening tools like the Cincinnati Pre-hospital Stroke Scale (CPSS) and Face Arm Speech Test (FAST) require specialized training for accurate administration.
- Limitations in current screening methods necessitate advanced diagnostic tools for acute settings.
Purpose of the Study:
- To propose and evaluate a novel multimodal deep learning approach for assessing suspected stroke patients.
- To leverage the Face Arm Speech Test (FAST) framework for developing an AI-driven stroke assessment tool.
- To enhance the accuracy and sensitivity of early stroke identification in emergency clinical settings.
Main Methods:
- Collected a multimodal dataset of emergency room patients performing FAST-based tests, including videos and audio recordings.
- Developed a novel deep learning model designed to process multi-modal data (action videos and speech audio).
- Compared the proposed multimodal model against six established action classification models (I3D, SlowFast, X3D, TPN, TimeSformer, MViT).
Main Results:
- The developed multimodal deep learning model demonstrated higher clinical value compared to existing action classification models.
- The multimodal model significantly outperformed its single-module variants, underscoring the importance of integrating diverse patient data.
- The model effectively processed video and audio data to assess symptoms like limb weakness, facial paresis, and speech disorders.
Conclusions:
- A multimodal deep learning model combined with the Face Arm Speech Test (FAST) offers a practical and powerful tool for early stroke identification.
- This approach significantly improves the accuracy and sensitivity of stroke assessment in emergency clinical settings.
- The findings highlight the potential of AI in revolutionizing acute stroke diagnostics and patient care.

