Related Experiment Video
Updated: Oct 2, 2026

Setup for the Quantitative Assessment of Motion and Muscle Activity During a Virtual Modified Box and Block Test
Published on: January 12, 2024
An Adaptive IMU-Visual Multimodal Fusion System for Real-Time Exercise Recognition and Movement Quality Assessment
Zhaoyang Gu1, Ruopeng Yang2, Yongqi Shi1
1Graduate School, National University of Defense Technology, Changsha 410073, China.
Abstract:
Exercise recognition and movement quality assessment remain challenging in supervised exercise training, particularly under viewpoint changes and self-occlusion. Vision-based methods provide spatial posture information but are susceptible to keypoint loss, whereas inertial sensing is less affected by occlusion but provides limited information about global posture geometry. This study presents a dual-stream prototype that combines a nine-axis inertial measurement unit (IMU) with vision-based pose estimation. A 1DCNN-LSTM branch models inertial dynamics, a custom keypoint temporal branch models normalized pose sequences, and a confidence-gated rule adjusts their contributions according to visual keypoint reliability. Owing to the absence of a public synchronized multi-view IMU-vision exercise dataset with the required protocol, we constructed IMV-Exercise, comprising 10 participants, three exercises, and 900 repetition-level samples with side-, front-, and posterior-view recordings. The system achieved 96.0% exercise recognition accuracy under leave-one-subject-out cross-validation. In a separate viewpoint-specific evaluation, the fused output achieved 91.2% action-window accuracy under posterior viewing. Across 50 online trials, the reported recognition accuracy was 96.0%, the mean end-to-end latency was 195 ms, and the recorded maximum was below 210 ms. These results establish feasibility within the studied cohort and exercises; broader generalization and feedback effectiveness require larger, independently controlled evaluations.