Related Experiment Video
Updated: Sep 26, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
Online GP-MPC Command Supervision for Robust Reinforcement Learning-Based Quadruped Locomotion
Seungyeon Lee1, Hyunseok Yang1
1Department of Mechanical Engineering, Yonsei University, Seoul 03722, Republic of Korea.
Abstract:
Reinforcement learning-based quadruped locomotion policies can exhibit command-tracking errors under terrain variations and unmodeled dynamics. This study proposes an online bounded Gaussian Process-enhanced model predictive control framework, termed Gaussian Process-Model Predictive Control-Reinforcement Learning(GP-MPC-RL), for command-level supervision of a pretrained locomotion policy. A frozen PPO policy generates the low-level locomotion behavior, while an acados-based MPC supervisor adjusts the velocity command using a nominal command-response model. An online Gaussian Process learns the one-step residual between the nominal prediction and measured robot response, and its uncertainty-weighted forward-velocity correction is incorporated into the MPC prediction. The framework was evaluated in Isaac Lab using a Unitree Go2 quadruped robot model over 20 paired rough-terrain trials at target velocities of 0.3, 0.5, and 0.7 m/s; GP-MPC-RL reduced the mean forward-velocity root mean square error (RMSE) relative to PPO by 38.3%, 21.3%, and 10.1%, respectively. Under a 5 kg payload, GP-MPC-RL reduced velocity RMSE by 29.0% relative to MPC-RL and reduced the 0-5 kg payload-induced degradation by 53.0% (p = 0.019). The average supervisor computation time was 0.112 ms. These results indicate that GP residual adaptation is particularly effective when the nominal command-response model becomes inaccurate, improving robustness without retraining the underlying locomotion policy.
