How Reward and the Elimination of Dependence Thereon Lead Evolution to High-Bandwidth Learning1

Solvi Arnold1, Reiji Suzuki2, Takaya Arita3

  • 1National Institute of Information and Communications Technology Universal Communication Research Institute solvi.arnold@nict.go.jp.

Artificial Life
|July 28, 2026
PubMed
Summary

This study introduces High-Bandwidth Learning (HBL), a novel AI approach inspired by biological learning efficiency. By evolving reward-driven learning systems, AI agents achieved significantly faster learning and greater autonomy, overcoming the limitations of traditional Reinforcement Learning (RL).

Related Concept Videos

Law of Effect01:06

Law of Effect

B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Evolution of New Traits in Microbes01:24

Evolution of New Traits in Microbes

Microorganisms evolve rapidly due to their large population sizes and short generation times, often exhibiting measurable changes within days under laboratory conditions. Natural selection acts on standing genetic variation, enabling the retention and amplification of beneficial traits that confer fitness advantages in changing environments.Adaptive Pigment Regulation in RhodobacterIn Rhodobacter, a genus of purple non-sulfur bacteria, light-harvesting pigments such as bacteriochlorophyll and...
Cognitive Learning01:21

Cognitive Learning

Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Purposive Learning01:22

Purposive Learning

E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...