Related Experiment Video
Updated: Jun 30, 2026

A Rehabilitation Program of Exoskeleton-assisted Body Weight-Supported Treadmill Training with Non-immersive Virtual Reality for Stroke Patients
Published on: May 16, 2025
Closed-Loop digital therapeutics empowered by deep reinforcement learning and wearable sensing for precision
Jiahao Dong1, Tao Chen2, Zhongyu Peng2
1First Clinical Medical College, Yunnan University of Traditional Chinese Medicine, Kunming, China.
Objective:
Static, one-size-fits-all protocols in postoperative orthopedic rehabilitation fail to adapt to individual recovery dynamics, potentially leading to suboptimal rehabilitation efficiency or an elevated risk of secondary injury. To address this critical gap, we propose and provide a simulation-based proof-of-concept validation for a novel closed-loop management system that deeply integrates patient-generated health data (PGHD) with deep reinforcement learning (DRL), offering a potential technical pathway for real-time personalized optimization of rehabilitation regimens and continuous prediction of long-term functional outcomes.
Methods:
A three-tier system architecture was constructed, comprising an intelligent sensing layer, an AI decision-making and prediction layer, and an interactive feedback layer. Through wearable inertial measurement units (IMUs) and surface electromyography (sEMG) devices, the system continuously collected multi-dimensional PGHD, including movement quality, training intensity, adherence, and pain feedback. These heterogeneous data were encoded into a comprehensive "patient state space" through a standardized feature engineering pipeline. A proximal policy optimization (PPO) algorithm was employed to train a DRL agent to learn the optimal policy for dynamically adjusting the next-cycle rehabilitation prescription (including exercise type, intensity, frequency, and progression pace). The agent aimed to maximize a hierarchical cumulative reward function that integrates short-term safety (ΔVAS pain score monitoring), mid-term adherence (training completion rate), and long-term functional improvement. Critically, a temporal convolutional network (TCN) prognostic module was deeply coupled with the DRL agent, providing prospective predictions of future functional recovery curves that inform the agent's long-term reward calculation, equipping the system with the capability of "making decisions based on predictions."
Results:
A proof-of-concept validation was conducted in a simulated environment for a post-operative anterior cruciate ligament reconstruction (ACLR) scenario. The trained DRL agent demonstrated the ability to generate differentiated rehabilitation strategies: in stratified analysis, it prescribed distinct progression paces for virtual patients with fast vs. slow recovery trajectories. Under this simulation setting, compared with a static conservative protocol, the DRL-driven strategy reduced the simulated time to "safe return to light activity" by an average of 15% and, compared with a static aggressive protocol, relatively decreased simulated "re-injury" events by 40%. The mean absolute errors (MAEs) of the TCN prognostic module for predicting functional scores at 2, 4, and 8 weeks into the future were 3.2, 4.8, and 6.5 points (on a 100-point scale), respectively, outperforming the ARIMA, LSTM, and GRU baseline models. An ablation study confirmed the TCN module's independent contribution, as its removal led to a relative increase in the simulated re-injury rate.
Conclusion:
This proof-of-concept study provides foundational evidence for the technical feasibility of a DRL-based closed-loop rehabilitation system. The proposed framework uniquely couples a wearable sensing layer with a symbiotic DRL-TCN architecture, demonstrating the potential to safely and dynamically personalize rehabilitation strategies in a simulated environment. These findings lay the groundwork for future prospective clinical trials, which are the necessary next step to validate safety, efficacy, and clinical utility in real-world settings.
