Related Experiment Video
Updated: Sep 21, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral Lagrangian
This study introduces a novel Separated Proportional-Integral Lagrangian (SPIL) algorithm to improve safety in reinforcement learning (RL). SPIL reduces oscillations and conservatism, enhancing policy reliability for real-world applications.
Area of Science:
- Artificial Intelligence
- Robotics
- Control Theory
Background:
- Real-world reinforcement learning (RL) demands robust safety guarantees under uncertainty.
- Existing chance-constrained RL methods often suffer from policy oscillations or suboptimal performance (overly conservative or unsafe).
Purpose of the Study:
- To develop an advanced algorithm that overcomes the limitations of current chance-constrained RL methods.
- To enhance the safety and reliability of RL policies in dynamic and uncertain environments.
Main Methods:
- Proposing a Separated Proportional-Integral Lagrangian (SPIL) algorithm, unifying proportional and integral control concepts.
- Employing an integral separation technique to manage integral values and stabilize learning.
- Utilizing model-based computation for the safe probability gradient to accelerate training.
Main Results:
- Demonstrated reduction in policy oscillations and conservatism in a car-following simulation.
- Successful application to a real-world mobile robot navigation task, ensuring obstacle avoidance despite unpredictable behaviors.
Conclusions:
- The SPIL algorithm offers a more stable and effective approach to chance-constrained reinforcement learning.
- SPIL enhances the practical applicability of RL in safety-critical real-world scenarios.
More Related Videos
06:45Design and Application of a Fault Detection Method Based on Adaptive Filters and Rotational Speed Estimation for an Electro-Hydrostatic Actuator
Published on: October 28, 2022
10:51An Experimental Platform to Study the Closed-loop Performance of Brain-machine Interfaces
Published on: March 10, 2011
Related Concept Videos
PD Controller: Design
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
PI Controller: Design
Phase-lead and Phase-lag Controllers
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Time and frequency -Domain Interpretation of PI Control
Acting as a low-pass filter, the PI controller slows the system's response and extends settling times. This requires...