Related Experiment Video
Updated: Aug 5, 2026

Studying Food Reward and Motivation in Humans
Published on: March 19, 2014
How Reward and the Elimination of Dependence Thereon Lead Evolution to High-Bandwidth Learning1
Solvi Arnold1, Reiji Suzuki2, Takaya Arita3
1National Institute of Information and Communications Technology Universal Communication Research Institute solvi.arnold@nict.go.jp.
This study introduces High-Bandwidth Learning (HBL), a novel AI approach inspired by biological learning efficiency. By evolving reward-driven learning systems, AI agents achieved significantly faster learning and greater autonomy, overcoming the limitations of traditional Reinforcement Learning (RL).
Area of Science:
- Artificial Intelligence
- Computational Neuroscience
- Machine Learning
Background:
- Current AI learning methods, particularly Reinforcement Learning (RL), suffer from low efficiency compared to biological learners, attributed to the 'Cost of Generality' and a 'Reward Bottleneck'.
- Biological intelligence efficiently processes rich sensory information, often independent of explicit reward signals, suggesting alternative learning mechanisms.
- Existing AI struggles to replicate this efficiency due to reliance on sparse, scalar reward feedback, limiting learning bandwidth.
Purpose of the Study:
- To propose and computationally explore High-Bandwidth Learning (HBL) as a biologically plausible pathway to overcome the Reward Bottleneck in AI.
- To investigate if reward-driven learning can act as a catalyst for the evolution of HBL, enabling AI to learn from diverse information streams.
- To demonstrate the scalability and effectiveness of evolved HBL agents in both simple and complex, high-dimensional tasks.
Main Methods:
- A population-based evolutionary approach was used, modeling reward-driven learning with RL (A2C) and incorporating a neuromodulatory connection weight update mechanism for integrating non-reward information.
- Experiments were conducted on a minimalistic 2D navigation task to assess initial learning speed and efficiency improvements.
- Scalability was tested on a high-dimensional quadruped locomotion task with randomized physical parameters.
Main Results:
- Evolved HBL agents demonstrated a 300-fold increase in learning speed on the navigation task compared to pure RL agents.
- Agents successfully eliminated dependence on reward, learning exclusively from non-reward information via local neuromodulation.
- In the quadruped task, evolved HBL agents significantly outperformed RL baselines in both learning speed and final performance, driven by reward-agnostic neuromodulation.
Conclusions:
- Reward-driven learning can serve as an evolutionary precursor to High-Bandwidth Learning, facilitating the gradual integration of non-reward information.
- HBL offers a promising, biologically inspired framework for developing more efficient, autonomous, and general artificial intelligence.
- The findings have implications for understanding the Baldwin effect, the evolution of autonomy, and the generality of human intelligence.
Related Concept Videos
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Evolution of New Traits in Microbes
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Purposive Learning