Related Experiment Video
Updated: Dec 6, 2025

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
9.0K
PORF-DDPG: Learning Personalized Autonomous Driving Behavior with Progressively Optimized Reward Function
Jie Chen1, Tao Wu1, Meiping Shi1
1The College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China.
Sensors (Basel, Switzerland)
|October 6, 2020
Summary
This study introduces a human-in-the-loop Deep Reinforcement Learning (DRL) algorithm for personalized autonomous driving. It uses a novel progressively optimized reward function (PORF) to improve safety and adaptability in dynamic scenarios.
Area of Science:
- Artificial Intelligence
- Robotics
- Machine Learning
Background:
- Deep Reinforcement Learning (DRL) shows promise for autonomous driving but struggles with complex, dynamic scenarios due to predefined reward functions.
- Ensuring safe and comfortable autonomous vehicle operation in real-world conditions remains a significant challenge for current DRL methods.
Purpose of the Study:
- To develop a human-in-the-loop DRL algorithm for learning personalized autonomous driving behaviors.
- To address the limitations of fixed reward functions in DRL for real-world driving scenarios.
- To enhance the safety and adaptability of autonomous vehicles in dynamic environments.
Main Methods:
- A novel progressively optimized reward function (PORF) learning model is proposed and integrated into the Deep Deterministic Policy Gradient (DDPG) framework, creating PORF-DDPG.
- The PORF model combines a pre-defined reward function with a Deep Neural Network (DNN) that learns driving intentions from human observers.
- The DNN-based reward component is progressively trained using front-view images and active human supervision.
Main Results:
- The proposed PORF-DDPG algorithm demonstrates online learning capabilities.
- The method shows adaptability to different environmental conditions.
- Experimental results indicate improved autonomous driving behavior learning compared to classic DRLs, especially in potentially hazardous situations.
Conclusions:
- The human-in-the-loop approach with a progressively optimized reward function enhances autonomous driving behavior learning.
- The PORF-DDPG method offers a promising solution for safe and personalized autonomous driving in complex dynamic scenarios.
- The algorithm's ability to learn from human supervision and adapt to environments suggests potential for real-world deployment.
Keywords:
autonomous drivingdeep reinforcement learningprogressive optimizationreward functionsequential framesMore Related Videos
Related Concept Videos
Observational Learning
697
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
697
Avoidance Learning and Learned Helplessness
2.3K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.3K
Purposive Learning
329
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
329
Behavior Modification
432
Behavioral approaches have often been criticized for ignoring mental processes and focusing solely on observable behavior. However, these approaches provide an optimistic perspective for individuals seeking to change their behaviors. Rather than concentrating on intrinsic personality traits, behavioral approaches suggest that even longstanding habits can be modified by changing the reward contingencies that maintain them.
A real-world application of operant conditioning principles is applied...
A real-world application of operant conditioning principles is applied...
432
PD Controller: Design
503
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
503
Associative Learning
948
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
948

