Multivideo Models for Classifying Hand Impairment After Stroke Using Egocentric Video.
Summary
Analyzing multiple videos of daily activities using a novel deep learning approach significantly improves the estimation of hand impairment in stroke survivors. This multi-video analysis offers a more accurate assessment than single-video methods for rehabilitation.
Area of Science:
- Biomedical Engineering
- Computer Vision
- Rehabilitation Science
Background:
- Current hand function assessments after stroke do not accurately reflect real-world performance.
- Wearable cameras capture hand function during activities of daily living (ADLs), but existing analysis methods focus on single tasks.
- There is a need for advanced analytical methods to process egocentric video data for comprehensive hand function assessment.
Purpose of the Study:
- To develop and evaluate a novel multi-video deep learning architecture for improved hand impairment estimation in stroke survivors.
- To compare the performance of multi-video analysis against single-video analysis using egocentric video data.
- To explore different fusion techniques for integrating information from multiple videos.
Main Methods:
- Developed single and multi-input video deep learning models using the SlowFast feature extractor.
- Utilized an egocentric video dataset of stroke survivors performing ADLs in a simulated home environment.
- Investigated late fusion (majority voting, fully-connected network) and intermediate fusion (concatenation, Markov chain) for multi-video architectures.
Main Results:
- The multi-video architecture using intermediate concatenation fusion achieved superior performance compared to single-video models.
- Multi-video models demonstrated significantly higher F1-scores (0.778 for cropped, 0.796 for full-frame inputs) than single-video models (0.696 and 0.708, respectively).
- Leave-One-Participant-Out-Cross-Validation confirmed the robustness and effectiveness of the multi-video approach.
Conclusions:
- Multi-video deep learning architectures provide a significant benefit for estimating hand impairment from egocentric video data post-stroke.
- The proposed approach represents a pioneering advancement in multi-video analysis for clinical applications.
- This methodology opens avenues for automating other complex, multi-observation clinical assessments.


