Related Experiment Video
Updated: Nov 19, 2025

The Modular Design and Production of an Intelligent Robot Based on a Closed-Loop Control Strategy
Published on: October 14, 2017
Learning-Based End-to-End Path Planning for Lunar Rovers with Safety Constraints
Xiaoqiang Yu1, Ping Wang2, Zexu Zhang1
1School of Astronautics, Harbin Institute of Technology, Harbin 150002, China.
This study introduces a novel learning-based path planning algorithm for lunar rovers, enhancing autonomous exploration safety. The developed deep reinforcement learning approach ensures safer navigation on complex lunar terrains.
Area of Science:
- Robotics
- Artificial Intelligence
- Planetary Science
Background:
- Autonomous exploration of the Moon requires robust path planning for lunar rovers.
- Existing algorithms may not adequately address safety constraints or real-world lunar terrain complexities.
Purpose of the Study:
- To propose a learning-based, end-to-end path planning algorithm for lunar rovers incorporating safety constraints.
- To enhance the safety and efficiency of autonomous lunar exploration missions.
Main Methods:
- Developed a Gazebo simulation environment with real lunar terrain data and a lunar rover simulator.
- Designed a deep reinforcement learning (DRL) algorithm, including state/action spaces, network architecture, and a reward function considering slip behavior.
- Employed proximal policy optimization (PPO) for training and curriculum learning to improve generalization across diverse lunar terrains and scales.
Main Results:
- The proposed DRL algorithm successfully achieved end-to-end path planning for lunar rovers in simulations.
- Paths generated by the algorithm demonstrated a higher safety guarantee compared to classical path planning methods.
- The use of curriculum learning improved the algorithm's generalization capabilities for varied lunar environments.
Conclusions:
- The learning-based end-to-end path planning algorithm offers a promising solution for safe and efficient lunar rover exploration.
- The integration of safety constraints and DRL, coupled with curriculum learning, significantly enhances navigation capabilities on extraterrestrial surfaces.
Related Concept Videos
Rolling Resistance: Problem Solving
Impact: Problem Solving
By designating the launch point as the origin and utilizing kinematic equations, the vertical component of the projectile's velocity at the point of impact is...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Circular Orbits and Critical Velocity for Satellites
Nicolaus Copernicus (1473-1543) first suggested that the Earth and all other planets orbit the Sun in...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Equation of Motion: General Plane motion - Problem Solving
The friction between the roller and the ground is characterized by two coefficients. The static friction coefficient is 0.15, while the kinetic friction coefficient is 0.1. These values are crucial in understanding the interaction between...

