Related Experiment Video
Updated: Jul 12, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
End-to-End Autonomous Navigation Based on Deep Reinforcement Learning with a Survival Penalty Function.
Shyr-Long Jeng1, Chienhsun Chiang2
1Department of Mechanical Engineering, Lunghwa University of Science and Technology, Taoyuan City 333326, Taiwan.
This study introduces a deep reinforcement learning approach for autonomous navigation in unknown dynamic environments. The method enhances robot survival and target achievement using a novel reward function, enabling collision-free path planning.
Area of Science:
- Robotics
- Artificial Intelligence
- Machine Learning
Background:
- Autonomous navigation in dynamic, map-less environments presents significant challenges.
- Traditional methods struggle with sparse rewards and complex obstacle avoidance.
Purpose of the Study:
- To propose an end-to-end deep reinforcement learning (DRL) approach for autonomous navigation.
- To enable nonholonomic wheeled mobile robots (WMRs) to navigate dynamic, map-less environments effectively.
- To address the sparse reward problem and ensure collision-free path planning.
Main Methods:
- Utilized two actor-critic (AC) frameworks: deep deterministic policy gradient (DDPG) and twin-delayed DDPG (TD3).
- Introduced a comprehensive reward function incorporating a survival penalty to guide the WMR towards its target.
- Connected consecutive episodes to increase cumulative penalties for obstacle scenarios, preventing training failure.
Main Results:
- Simulations in various scenarios (obstacle-free, parking lot, intersections, multiple obstacles) demonstrated the method's efficiency and safety.
- The TD3 algorithm showed faster convergence and greater stability during training compared to DDPG.
- TD3 achieved a higher task execution success rate during evaluation.
Conclusions:
- The proposed DRL approach with a survival penalty function effectively enables autonomous navigation in challenging environments.
- The TD3 algorithm offers superior performance in terms of training efficiency and navigation success rate over DDPG.
- This method provides a robust solution for collision-free path planning for WMRs.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Rolling Resistance: Problem Solving
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Survival Tree
Building a Survival Tree
Constructing a...
Hydraulic Jump: Problem Solving
Observational Learning

