Related Experiment Video
Updated: Aug 28, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
Multimodal Inertial-Visual Sensor Fusion over Evolutionary Deep Temporal Modeling for Humanoid Movement Recognition:
Mohammad Shorfuzzaman1, Muhammad Hanzla2, Bayan Alabdullah3
1Department of Software Engineering, College of Engineering and Advanced Computing, Alfaisal University, Riyadh 11533, Saudi Arabia.
Abstract:
Wearable inertial sensing and markerless vision are increasingly integrated to enable objective assessment of locomotor and postural function for sports telerehabilitation, intelligent physiotherapy, and athlete performance monitoring. Before such multimodal systems can be translated to clinical practice, their fusion, optimization, and temporal modeling strategies require validation under controlled conditions with reliable ground truth. This study presents a unified multimodal framework that hierarchically integrates inertial measurement unit (IMU) signals and RGB visual information through kernelized representation learning, adaptive multimodal fusion, evolutionary feature optimization, and deep temporal classification. The IMU branch employs Kernelized Extreme Learning Machine (KELM) denoising, Kernelized Canonical Correlation Fusion (KCCF), entropy-guided adaptive windowing, and complementary time-series descriptors (MINIROCKET, TS-CHIEF, and r-STSF). Concurrently, the RGB branch combines anisotropic diffusion filtering, HRNet-based silhouette extraction, DensePose R-CNN, Mesh Graphormer, and Multi-Model Pose-Flow Fusion (MPFF) to learn robust visual representations. Both modalities are integrated through Weighted Canonical Feature Fusion (WCFF) and optimized using a Genetic Algorithm for feature selection and adaptive modality weighting before temporal modeling with cluster-based alignment, Gaussian Process Sequence Modeling, and DeepConvLSTM. As the selected benchmarks do not provide complete inertial recordings, the inertial modality is established according to the adopted experimental protocol to support multimodal fusion analysis. Under 5-fold subject-independent cross-validation, the framework achieves accuracies of 86.56 ± 0.31% on SoccerDiffusion and 88.04 ± 0.25% on HumanoidRobotPose. Although evaluated on humanoid robotic benchmarks, the proposed framework provides a methodological basis for future wearable-enabled clinical movement assessment, remote rehabilitation, and athlete monitoring, while validation on synchronized human inertial-visual datasets remains an important direction for future research.