Related Experiment Video
Updated: Sep 24, 2025

Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
Discovering diverse solutions in deep reinforcement learning by maximizing state-action-based mutual information.
Takayuki Osa1, Voot Tangkaratt2, Masashi Sugiyama3
1Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu, Kita-kyushu, 808-0135, Fukuoka, Japan; RIKEN Center for Advanced Intelligence Project, 1-4-1 Nihonbashi, Chuo-ku, 103-0027, Tokyo, Japan.
This study introduces a new reinforcement learning method to generate diverse solutions for tasks, improving few-shot adaptation. The approach avoids bias issues found in prior techniques.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Reinforcement learning (RL) typically produces single solutions, despite task diversity.
- Learning diverse solutions enhances few-shot adaptation capabilities.
- Existing methods using mutual information for diversity suffer from gradient estimator bias.
Purpose of the Study:
- To develop a novel reinforcement learning method for learning diverse solutions without gradient bias.
- To enable learning an infinite set of diverse solutions using continuous latent variables.
- To improve few-shot adaptation compared to existing approaches.
Main Methods:
- A policy conditioned on latent variables is trained by directly maximizing the variational lower bound of mutual information.
- This method avoids using mutual information as unsupervised rewards, mitigating bias.
- Experiments conducted on robot locomotion tasks.
Main Results:
- The proposed method successfully learns an infinite set of diverse solutions using continuous latent variables.
- Demonstrated superior few-shot adaptation performance compared to existing methods.
- Effectively addresses the bias problem inherent in prior techniques.
Conclusions:
- The novel method effectively learns diverse solutions in reinforcement learning.
- Continuous latent variables facilitate learning an infinite set of diverse policies.
- The approach offers improved robustness and adaptation capabilities for complex tasks.
Related Concept Videos
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Multi-input and Multi-variable systems
In the absence...
Principle of Moments: Problem Solving
One such scenario involves a pole placed in a three-dimensional system with a cable attached. When a tension is applied to the cable, the moment about the z-axis passing through...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

