Related Experiment Video
Updated: Sep 12, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Optimal control under safety constraints and disturbances: a multi-step, off-policy adaptive dynamic programming
Jun Ye1,2, Xiaowei Zhao2, Yougang Bian1
1State Key Laboratory of Advanced Design and Manufacturing Technology for Vehicle, College of Mechanical and Vehicle Engineering, Hunan University, Changsha, 410012 China.
This study presents a novel adaptive dynamic programming method for optimal control under safety constraints. It effectively balances performance and safety using dual critics and an actor-critic-disturbance framework.
Area of Science:
- Robotics and Control Systems
- Artificial Intelligence
- Optimization Theory
Background:
- Optimal control problems often involve complex dynamics and safety requirements.
- Traditional methods struggle with disturbances and ensuring constraint satisfaction.
- Accurate performance estimation is crucial but challenging under uncertainty.
Purpose of the Study:
- To introduce a robust multi-step, off-policy adaptive dynamic programming (ADP) approach for optimal control.
- To address challenges posed by external disturbances and critical safety constraints.
- To enhance the accuracy of performance function estimation and ensure a safety-performance trade-off.
Main Methods:
- Developed model-free and model-based variants of the off-policy adaptive dynamic programming.
- Employed interleaved training and prior models to mitigate underestimation in policy evaluation.
- Utilized dual critic neural networks to counteract terminal performance function underestimation.
- Transformed policy improvement into a constrained optimization task with a safety function.
- Designed an actor-critic-disturbance framework for handling safety constraints within a zero-sum game.
Main Results:
- The proposed method demonstrates effectiveness in solving optimal control problems with disturbances and safety constraints.
- Theoretical analysis confirms the convergence properties of the adaptive dynamic programming approach.
- Simulation results and practical experiments validate the method's performance and safety.
Conclusions:
- The multi-step, off-policy adaptive dynamic programming approach offers a viable solution for complex optimal control scenarios.
- The integration of dual critics and the actor-critic-disturbance framework enhances safety and performance.
- The method shows significant promise for real-world applications requiring safe and efficient control.
Related Concept Videos
Time-Domain Interpretation of PD Control
Consider the example of control of motor torque. Initially, a positive...
PD Controller: Design
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Control Systems
At the heart...
Time and frequency -Domain Interpretation of PI Control
Acting as a low-pass filter, the PI controller slows the system's response and extends settling times. This requires...
PI Controller: Design
Statically Indeterminate Problem Solving

