Related Experiment Video
Updated: Nov 10, 2025

Tactile Vibrating Toolkit and Driving Simulation Platform for Driving-Related Research
Published on: December 18, 2020
Weakly Supervised Reinforcement Learning for Autonomous Highway Driving via Virtual Safety Cages
Sampo Kuutti1, Richard Bowden2, Saber Fallah1
1Connected and Autonomous Vehicles Lab, University of Surrey, Guildford GU2 7XH, UK.
This study introduces rule-based safety cages to enhance reinforcement learning for autonomous vehicle control. These cages improve safety, training speed, and policy performance, even with suboptimal parameters.
Area of Science:
- Autonomous Systems
- Artificial Intelligence
- Control Theory
Background:
- Neural networks and reinforcement learning (RL) are increasingly used for autonomous vehicle control.
- The 'black box' nature of RL policies hinders trust and deployment in safety-critical applications.
- Lack of interpretability in RL control policies is a major barrier for autonomous vehicle adoption.
Purpose of the Study:
- To develop a reinforcement learning approach for autonomous vehicle longitudinal control.
- To enhance safety and interpretability using rule-based safety cages as weak supervision.
- To improve training convergence and the safety of the final control policy.
Main Methods:
- Implemented a reinforcement learning agent for autonomous vehicle longitudinal control.
- Integrated rule-based safety cages to provide weak supervision and enhance safety.
- Compared RL models with and without safety cages, and with optimal vs. constrained parameters.
Main Results:
- Weak supervision from safety cages consistently improved exploration safety, convergence speed, and model performance.
- Safety cages enabled learning a safe driving policy even when RL alone failed with constrained parameters.
- The rule-based supervisory controller offers interpretability, facilitating traditional validation and verification.
Conclusions:
- Rule-based safety cages are effective for enhancing RL-based autonomous vehicle control.
- This approach addresses the interpretability and safety challenges of deploying RL in autonomous vehicles.
- The method allows for safe policy learning even under suboptimal or constrained model conditions.
More Related Videos
06:28A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
11:12Driving Simulation in the Clinic: Testing Visual Exploratory Behavior in Daily Life Activities in Patients with Visual Field Defects
Published on: September 18, 2012
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Once a behavior is learned,...
Rolling Resistance: Problem Solving