Related Experiment Video
Updated: Jan 29, 2026

Experimental Investigation of the Hierarchical Control in DC Microgrids Using a Real-time Simulator
Published on: February 14, 2025
Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning
This study introduces a deep reinforcement learning (DRL) algorithm to enhance data efficiency by integrating unsupervised learning, intrinsic motivation, and hierarchical reinforcement learning (HRL). The novel approach improves agent exploration and skill acquisition, outperforming baselines in Atari games.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Deep Reinforcement Learning (DRL) faces challenges with data inefficiency due to underutilization of acquired data and suboptimal exploration strategies.
- Effective utilization of experiences and improved exploration are critical for advancing DRL capabilities.
Purpose of the Study:
- To propose a DRL algorithm that enhances data efficiency by leveraging unrewarded experiences and refining exploration strategies.
- To combine unsupervised auxiliary tasks, intrinsic motivation, and hierarchical reinforcement learning (HRL) for improved DRL performance.
Main Methods:
- Developed a DRL algorithm based on a hierarchical reinforcement learning (HRL) architecture with a metacontroller and subcontroller.
- Integrated intrinsic motivation for the subcontroller, guided by the metacontroller, to learn environment control and reusable skills.
- Reinterpreted auxiliary tasks as skills learned through intrinsic rewards for temporally extended exploration.
Main Results:
- The proposed DRL algorithm demonstrated superior performance compared to baseline methods across several Atari 2600 games.
- Achieved significant performance gains in Montezuma's Revenge, a notoriously difficult game requiring sparse data utilization.
- Confirmed the critical role of intrinsic rewards in performance enhancement, with learned representations contributing substantially to the benefits.
Conclusions:
- The novel DRL approach effectively addresses data inefficiency by improving experience utilization and exploration.
- Intrinsic rewards and learned representations are key components for advancing DRL performance, particularly in complex environments with sparse rewards.
- The method shows promise for developing more capable and data-efficient DRL agents applicable to challenging tasks.
Related Concept Videos
Secondary Motives: Power Motivation and Achievement Motivation
Power motivation is characterized by the desire to influence, control, or have an impact on others. It is shaped by an individual's experiences, social environment, and cultural context. People with high power motivation are...
Secondary Motives: Affiliation Motivation and Aggression Motivation
Intrinsically Disordered Proteins
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Motivational Bias
Motivational Cycle
The cycle begins with a need. This need can arise from various conditions, such as hunger, thirst, or temperature changes. For instance, when an individual feels cold, their body...

