Related Experiment Video
Updated: Oct 8, 2025

Author Spotlight: Investigating the Effects of Mind-Body-Movement Practices on Brain Function
Published on: January 26, 2024
End-to-End Autonomous Exploration with Deep Reinforcement Learning and Intrinsic Motivation
Xiaogang Ruan1,2, Peng Li1,2, Xiaoqing Zhu1,2
1Faculty of Information Technology, Beijing University of Technology, Beijing, China.
Abstract:
Developing artificial intelligence (AI) agents is challenging for efficient exploration in visually rich and complex environments. In this study, we formulate the exploration question as a reinforcement learning problem and rely on intrinsic motivation to guide exploration behavior. Such intrinsic motivation is driven by curiosity and is calculated based on episode memory. To distribute the intrinsic motivation, we use a count-based method and temporal distance to generate it synchronously. We tested our approach in 3D maze-like environments and validated its performance in exploration tasks through extensive experiments. The experimental results show that our agent can learn exploration ability from raw sensory input and accomplish autonomous exploration across different mazes. In addition, the learned policy is not biased by stochastic objects. We also analyze the effects of different training methods and driving forces on exploration policy.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Ryan and Deci's Self-Determination Theory
Autonomy is the need to feel in control of one's actions and decisions. For example, a student who chooses their research topic is likely to be more engaged and motivated than one who...
Purposive Learning
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Reinforcement Schedules
Once a behavior is learned,...

