Related Experiment Video
Updated: May 20, 2026

06:57
Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
An information-theoretic approach to curiosity-driven reinforcement learning
1Information and Computer Sciences, University of Hawaii at Mānoa, Honolulu, HI 96822, USA. sstill@hawaii.edu
Theory in Biosciences = Theorie in Den Biowissenschaften
|July 14, 2012
Summary
This study reframes exploration in reinforcement learning using information theory. It reveals Boltzmann exploration is optimal and proposes curiosity-driven learning maximizes predictive power for better exploration-exploitation trade-offs.
Area of Science:
- Artificial Intelligence
- Information Theory
- Machine Learning
Background:
- Reinforcement learning (RL) relies on exploration strategies to balance reward acquisition and learning.
- Boltzmann-style exploration is a common method, but its theoretical optimality is re-examined.
- Curiosity-driven learning aims to enhance agent engagement and knowledge acquisition.
Purpose of the Study:
- To provide an information-theoretic perspective on exploration in reinforcement learning.
- To demonstrate the optimality of Boltzmann-style exploration concerning expected return and policy coding cost.
- To introduce a novel framework for curiosity-driven learning that integrates predictive power maximization.
Main Methods:
- Information-theoretic analysis of exploration strategies.
- Derivation of optimal policies balancing expected return and information gain.
- Formulation of a new exploration-exploitation trade-off based on predictive power.
Main Results:
- Boltzmann-style exploration is shown to be information-theoretically optimal.
- A novel exploration bonus is introduced, enhancing curiosity-driven learning.
- The proposed exploration-exploitation trade-off is inherent in optimal deterministic policies.
Conclusions:
- Exploration in reinforcement learning can be viewed as optimizing information gain.
- Curiosity-driven learning enhances agent performance by maximizing predictive power.
- This approach offers a more principled understanding of exploration beyond random action selection.
Related Concept Videos
Cognitive Learning
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Purposive Learning
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...
Avoidance Learning and Learned Helplessness
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

