Related Experiment Video
Updated: Oct 2, 2026

An Inertial Measurement Unit Based Method to Estimate Hip and Knee Joint Kinematics in Team Sport Athletes on the Field
Published on: May 26, 2020
Reliability and Validity of AI-Based Pose Estimation Algorithms for Assessing Lower-Limb Flexibility and Joint Range
Hinata Okuno1, Tomoya Ishida2,3, Yuta Koshino2,3
1Graduate School of Health Sciences, Hokkaido University, Kita-12 Nishi-5, Kita-ku, Sapporo 060-0812, Hokkaido, Japan.
Abstract:
Muscle flexibility and joint range of motion (ROM) evaluation are essential for injury prevention and rehabilitation; however, conventional manual methods are limited in large-scale screening and self-monitoring because of examiner dependency and time constraints. This study evaluated the reliability and validity of AI-based pose estimation for lower-limb flexibility and ROM tests compared with conventional manual measurements. Twenty healthy volunteers (40 lower limbs; age 24.4 ± 2.6 years) underwent three flexibility tests-the Active Knee Extension Test (AKET), Modified Thomas Test (MTT), and Weight-Bearing Lunge Test (WBLT). The tests were assessed by a physiotherapist and three pose-estimation models (MediaPipe, OpenPose, and HRNet). Intra-rater reliability was evaluated with ICC(3,1). Concurrent validity and agreement were assessed using linear regression and Bland-Altman analysis. Intra-rater reliability was good to excellent for all methods (ICC ≥ 0.87), and concurrent validity was high (r2 ≥ 0.80), except for MTT using the HRNet (r2 = 0.73). Fixed biases were observed during the flexibility test. The limits of agreement ranged from ±8.04-10.87 cm for the AKET and MTT, and ±3.13-4.42 cm for WBLT. The WBLT showed a significant systematic bias, with an underestimation of -6.25° to -7.17° across all models. In young, healthy volunteers under standardized laboratory conditions, pose estimation models showed robust reliability and concurrent validity as objective tools for lower-limb flexibility assessment. However, systematic biases warrant caution regarding the interchangeability between AI-based measurements and conventional methods.