Related Experiment Video
Updated: Dec 3, 2025

Tactile Vibrating Toolkit and Driving Simulation Platform for Driving-Related Research
Published on: December 18, 2020
Policy-Gradient and Actor-Critic Based State Representation Learning for Safe Driving of Autonomous Vehicles
Abhishek Gupta1, Ahmed Shaharyar Khwaja1, Alagan Anpalagan1
1Department of Electrical, Computer and Biomedical Engineering, Ryerson University, Toronto, ON M5B2K3, Canada.
This study introduces a novel environment perception framework for autonomous driving using state representation learning (SRL). The method enhances learning for complex driving scenarios without explicit datasets, enabling safer navigation.
Area of Science:
- Robotics
- Artificial Intelligence
- Computer Vision
Background:
- Autonomous driving systems require robust environment perception for safe navigation.
- Existing methods often rely on extensive labeled datasets and may struggle with complex, dynamic environments.
- Q-learning based approaches offer efficiency but can be limited in handling stochastic policies.
Purpose of the Study:
- To develop an advanced environment perception framework for autonomous vehicles using state representation learning (SRL).
- To improve autonomous driving capabilities by integrating learning loss considerations into policy gradient methods.
- To achieve uninterrupted and safe driving over extended distances in complex environments.
Main Methods:
- The framework combines Variational Autoencoder (VAE) for representation learning with Deep Deterministic Policy Gradient (DDPG) and Soft Actor-Critic (SAC) for policy optimization.
- It incorporates learning loss within both deterministic and stochastic policy gradients.
- A reward-penalty system is implemented to guide learning, with positive rewards for favorable actions and negative rewards for unfavorable ones.
Main Results:
- Simulations on the DonKey simulator demonstrated the framework's effectiveness in learning complex environmental interactions.
- Analysis of policy loss, value loss, reward function, and cumulative reward showed successful learning progression for both VAE+DDPG and VAE+SAC configurations.
- The proposed method achieved uninterrupted driving without deviating from the track for significant distances.
Conclusions:
- The proposed state representation learning framework offers a promising approach for enhancing autonomous driving perception.
- Integrating learning loss into policy gradients and utilizing VAE, DDPG, and SAC enables robust learning from complex environmental interactions.
- The reward-penalty system effectively guides the autonomous vehicle towards safer and more sustained driving performance.
More Related Videos
Related Concept Videos
State Space Representation
Consider an RLC circuit, a...
Observational Learning
Controller Configurations
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Rolling Resistance: Problem Solving
Automatic Processing and Automatic Social Behavior
Hierarchy of Motor Control

