Related Experiment Video
Updated: Jul 25, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.0K
Adaptive Safe Reinforcement Learning With Full-State Constraints and Constrained Adaptation for Autonomous Vehicles
IEEE Transactions on Cybernetics
|June 26, 2023
Summary
This study introduces an adaptive safe reinforcement learning (RL) algorithm for autonomous vehicles, ensuring all state variables remain within safe limits during learning. The method enhances control performance and safety under uncertainty.
Area of Science:
- Autonomous Systems
- Machine Learning
- Control Theory
Background:
- Safety-critical autonomous vehicles require state variables to remain within defined regions during learning.
- Existing reinforcement learning (RL) methods may not guarantee safety constraints throughout the learning process.
Purpose of the Study:
- To develop an adaptive safe reinforcement learning (RL) algorithm for autonomous vehicles.
- To ensure full-state variables are constrained within the safety region during the entire learning process.
- To enhance control performance and safety assurance in autonomous vehicle systems.
Main Methods:
- An adaptive safe RL algorithm integrating an optimized backstepping technique and asymmetric barrier Lyapunov function (BLF) methodology.
- Decomposition of subsystem control and value function derivatives with BLF-related terms and independent learning components.
- A constrained adaptation algorithm with a projection operator to manage safety-optimization conflicts.
Main Results:
- The proposed algorithm ensures full-state variables remain within the safety region during learning.
- Demonstrated superior performance in simulations for autonomous vehicle motion control compared to existing methods.
- Verified improved convergence and reduced variance, particularly under uncertain conditions.
Conclusions:
- The adaptive safe RL algorithm effectively optimizes system control while guaranteeing state variable constraints.
- The methodology provides doubly assured safety performance through constrained adaptation.
- The approach is validated for enhancing the safety and performance of autonomous vehicle control systems.
Related Concept Videos
Controller Configurations
128
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
128
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K
Constraints and Statical Determinacy
640
In structural engineering, the equilibrium of a system is not only determined by its equations of equilibrium but also with the help of constraints. Constraints refer to restrictions on the motion of a system. The proper combinations of constraints can minimize the total number of constraints needed to maintain a system in mechanical equilibrium. When this happens, the system is said to be statically determinate. For such systems, the unknown reaction supports can be estimated using equilibrium...
640
PD Controller: Design
291
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
291
Rolling Resistance: Problem Solving
385
Rolling resistance, also known as rolling friction, is the force that resists the motion of a rolling object, such as a wheel, tire, or ball, when it moves over a surface. It is caused by the deformation of the object and the surface in contact with each other, as well as other factors like internal friction, hysteresis, and energy losses within the materials. Rolling resistance opposes the object's motion, requiring additional energy to overcome it and maintain movement. In practical...
385
Observational Learning
222
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
222

