Related Experiment Video
Updated: Jun 21, 2025

07:15
Tactile Vibrating Toolkit and Driving Simulation Platform for Driving-Related Research
Published on: December 18, 2020
4.5K
Towards Robust Decision-Making for Autonomous Highway Driving Based on Safe Reinforcement Learning.
Rui Zhao1, Ziguo Chen1, Yuze Fan1
1College of Automotive Engineering, Jilin University, Changchun 130025, China.
Sensors (Basel, Switzerland)
|July 13, 2024
Summary
This study introduces a new framework for safe autonomous driving using Replay Buffer Constrained Policy Optimization (RECPO). The method enhances reinforcement learning policies for robust highway driving, achieving zero collisions.
Area of Science:
- Artificial Intelligence
- Robotics
- Autonomous Systems
Background:
- Reinforcement Learning (RL) is effective for autonomous driving but struggles with robust safety in diverse data, especially long-tail scenarios.
- Ensuring safety and considering data distribution variations are critical challenges for RL-based decision-making in autonomous vehicles.
Purpose of the Study:
- To present a novel framework for highway autonomous driving that prioritizes both safety and robustness.
- To develop a method that updates RL strategies to maximize rewards while adhering to safety constraints.
Main Methods:
- Introduced Replay Buffer Constrained Policy Optimization (RECPO) to update RL strategies within safety constraints.
- Utilized importance sampling and a Replay buffer for data reutilization, mitigating catastrophic forgetting.
- Formulated the problem as a Constrained Markov Decision Process (CMDP) for policy optimization.
Main Results:
- The RECPO framework demonstrated significantly enhanced model convergence speed, safety, and decision-making stability in CARLA simulations.
- Achieved a zero-collision rate in highway autonomous driving scenarios.
- Outperformed traditional CPO, Deep Deterministic Policy Gradient (DDPG), and Intelligent Driver Model + MOBIL (IDM + MOBIL) methods.
Conclusions:
- The proposed RECPO method offers a robust and safe approach for training autonomous driving policies.
- The framework effectively addresses safety challenges in RL for autonomous driving, particularly in complex highway environments.
Keywords:
autonomous drivingcatastrophic forgettingconstrained policy optimizationdeep reinforcement learningimportance samplingMore Related Videos
Related Concept Videos
Decision Making: Traditional Method
4.0K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.0K
Reinforcement Schedules
140
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
140
PD Controller: Design
215
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
215
Decision Making: P-value Method
5.3K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.3K
Decision Making
106
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Automatic decision-making is fast, intuitive, and relies on gut feelings...
106
The Anchoring-and-Adjustment Heuristic
7.2K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.2K

