Differentiating post-stroke patients from healthy individuals via vision-based skeleton-optical fusion
Xiao Han1,2, Ziyan Wang3,4, Liping Li5
1Institute of Traditional Chinese Medicine Literature, Nanjing University of Chinese Medicine, No.138, Xianlin Road, Nanjing, 210046, China. x.han@njucm.edu.cn.
Background:
At present, the analysis of abnormal gait in post-stroke patients predominantly relies on wearable devices. However, with the advancements in computer vision technology, the integration of deep learning algorithms has introduced new possibilities for research. In particular, multi-modal fusion technology can effectively combine various modalities obtained through vision-based approaches, enabling more comprehensive and accurate representation of abnormal gait information in post-stroke patients.
Methods:
The study recruited 70 post-stroke patients and 70 healthy individuals to capture video recordings of their gait. We used Human Pose Estimation (HPE) to extract skeleton points from each frame and computed the optical flow information of these points and the corresponding angular variations of the lower limbs. Additionally, depth space features were extracted using ResNet-50 and subsequently integrated. For classification, a Long Short-Term Memory (LSTM) network was employed to analyze the fused features.
Results:
To evaluate the effectiveness of the feature extraction method, we tested it on both an open dataset and a self-collected clinical dataset, comparing it with CNN-RNN and Vision Transformer (ViT). The results from the LSTM network, after inputting the fused features, demonstrated optimal performance with 2 layers and 128 hidden units, achieving accuracies of 0.8794±0.0447 and 0.8778±0.0347, respectively.
Conclusion:
It was found that optical flow information calculated based on skeleton points, combined with variations in knee flexion and ankle dorsiflexion angles, improved the interpretability of the analytical framework. This improvement enables clinicians to gain a clearer understanding of the model's decision-making process, thereby increasing their confidence in its outputs. By employing a multi-modal fusion approach, information from different modalities is integrated, which not only broadens the analytical perspectives but also facilitates clinicians' deeper insights into the patient's gait characteristics.
More Related Videos
07:24Using Eye-tracking to Assess the Relative Importance of Visual and Vestibular Input to Subcortical Motion Processing in the Roll Plane
Published on: August 22, 2025
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
