Related Experiment Videos
Value-Regularized Reinforcement Learning for Model Predictive Control of Autonomous Mobile Robots Under Stochastic
Changyuan Yu1, Weiguo Zhang1, Qi Li1
1Xi'an Institute of Applied Optics, Xi'an 710065, China.
Sensors (Basel, Switzerland)
|July 28, 2026
Summary
Value-regularized reinforcement learning-based model predictive control (VR-RLMPC) enhances robot control. This method improves performance and reduces degradation under model mismatch, outperforming existing baselines.
Area of Science:
- Robotics
- Control Systems
- Machine Learning
Background:
- Autonomous mobile robots require robust control for reliable operation.
- Model predictive control (MPC) and reinforcement learning (RL) are key techniques, but face challenges with process disturbances and model mismatch.
- Existing RL-MPC methods often rely on nominal parameters, limiting performance in real-world scenarios.
Purpose of the Study:
- To introduce a novel control strategy, value-regularized RL-MPC (VR-RLMPC), to improve the robustness and performance of autonomous mobile robots.
- To enhance the RL-MPC framework by incorporating isotropic state regularization into the stage cost.
- To evaluate the effectiveness of VR-RLMPC against traditional RL-MPC and other baselines under various conditions, including model mismatch.
Main Methods:
- Introduced isotropic state regularization into the RL-MPC stage cost to create VR-RLMPC.
- Utilized a Lyapunov-drift analysis to establish theoretical performance bounds under specific assumptions.
- Conducted simulations on linear and nonlinear nonholonomic vehicle systems to compare VR-RLMPC with RL-MPC and stochastic temporal-difference baselines.
Main Results:
- VR-RLMPC demonstrated approximately 25% lower settling-step indices compared to RL-MPC across nominal and model-mismatch scenarios.
- The worst-case performance degradation rate for VR-RLMPC was significantly lower (2.8%) than conventional MPC baselines (25-49%) under model mismatch.
- VR-RLMPC showed improved cost performance with increasing task difficulty and outperformed a stochastic temporal-difference baseline in nominal settings.
Conclusions:
- VR-RLMPC offers a significant improvement in control robustness and performance for autonomous mobile robots, particularly under model mismatch.
- The proposed method effectively enhances RL-MPC without increasing computational complexity or online optimization requirements.
- VR-RLMPC represents a promising advancement for reliable autonomous system control in dynamic environments.
Related Concept Videos
Vector Functions and Motion: Problem Solving
Accurate position tracking is fundamental to the safe and effective operation of unmanned aerial vehicles (UAVs), particularly during precision maneuvers near complex structures. In this scenario, a drone is programmed to perform a high-precision inspection of a vertical structure, starting at position ((x, y, z) = (3, 0, 0)), with an initial velocity oriented in the positive z-direction. The trajectory of the drone is governed by a time-dependent acceleration function a(t), which is predefined...
Time-Domain Interpretation of PD Control
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
Control Systems
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
PD Controller: Design
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Rolling Resistance: Problem Solving
Rolling resistance, also known as rolling friction, is the force that resists the motion of a rolling object, such as a wheel, tire, or ball, when it moves over a surface. It is caused by the deformation of the object and the surface in contact with each other, as well as other factors like internal friction, hysteresis, and energy losses within the materials. Rolling resistance opposes the object's motion, requiring additional energy to overcome it and maintain movement. In practical...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...