Related Experiment Video
Updated: Oct 10, 2026

Video Movement Analysis Using Smartphones (ViMAS): A Pilot Study
Published on: March 14, 2017
Clinical knowledge-guided attention for gait-based adult spinal deformity assessment from monocular video
Kaixu Chen1, Tomoyuki Asada2, Kousei Miura2
1University of Tsukuba, Center for Computational Sciences, Tsukuba, Japan.
Purpose:
Adult spinal deformity (ASD) is associated with characteristic gait abnormalities, yet routine diagnosis relies mainly on radiographic assessment and physical examination, which can be time-consuming and may not reflect functional impairment during walking. Existing video-based gait models are often purely data-driven and difficult to interpret clinically. This study aims to develop an explainable, noninvasive ASD screening framework that integrates orthopedic clinical knowledge to guide model attention toward diagnostically meaningful gait cues.
Approach:
This paper proposes an anatomy-guided attention framework for ASD diagnosis from monocular gait videos. Human pose is extracted from videos and represented as spatiotemporal joint sequences. Orthopedic specialists provide clinical experience about anatomically important joints and motion patterns, which is encoded as clinician-informed attention maps. We then fuse these attention maps with the input videos to guide the classification model to focus on disease-relevant regions. The attention-guided maps are integrated with a three-dimensional (3D) CNN-based spatiotemporal backbone, and model visualization is used to assess interpretability.
Results:
Experiments on a gait dataset with 81 patients show that the proposed approach consistently outperforms conventional baselines. The proposed framework was evaluated using monocular gait videos from patients with adult spinal deformity, dropped head syndrome (DHS), lumbar canal stenosis (LCS), and hip osteoarthritis (HipOA). The proposed method achieves 71.35% accuracy, 75.51% precision, and 71.12% F1-score. In contrast, the model without attention achieves 62.09% accuracy, 64.55% precision, and 60.13% F1-score, corresponding to improvements of 9.26, 10.96, and 10.99 percentage points, respectively. Visualization analyses further show that anatomy-guided attention yields clearer and more stable focus on clinically meaningful body regions, producing explanations that better align with expert expectations.
Conclusions:
Incorporating clinician-derived priors as attention guidance improves both diagnostic accuracy and interpretability for gait-based ASD recognition from monocular videos. The proposed framework supports accessible, radiation-free screening and provides clinically informative visual evidence, facilitating translation of gait analysis models to real-world orthopedic practice.
