Related Experiment Video
Updated: Jan 5, 2026

An Open-Source Virtual Reality System for the Measurement of Spatial Learning in Head-Restrained Mice
Published on: March 3, 2023
Computational noise in reward-guided learning drives behavioral variability in volatile environments
Charles Findling1,2, Vasilisa Skvortsova1, Rémi Dromnelle1,3
1Laboratoire de Neurosciences Cognitives et Computationnelles, Inserm U960, Département d'Études Cognitives, École Normale Supérieure, PSL University, Paris, France.
Human decisions in changing environments are often irrational, not due to exploration, but computational noise in learning action values. This learning noise, impacting reward-guided learning, explains most behavioral variability.
Area of Science:
- Neuroscience
- Cognitive Science
- Computational Psychiatry
Background:
- Humans often make suboptimal choices in volatile environments, deviating from maximizing expected value.
- These 'non-greedy' decisions are traditionally viewed as information-seeking behavior or exploration.
Purpose of the Study:
- To investigate the computational basis of non-greedy decisions in volatile environments.
- To determine if behavioral variability stems from exploration or noise in value learning.
Main Methods:
- Utilized reinforcement learning models to simulate behavior.
- Collected multimodal neurophysiological data, including blood oxygen level-dependent (BOLD) responses and pupillary dilation.
- Analyzed trial-to-trial variability in sequential learning steps.
Main Results:
- The majority of non-greedy decisions were attributed to computational noise in action value learning.
- Trial-to-trial learning variability predicted behavior.
- BOLD responses in the dorsal anterior cingulate cortex and pupillary dilation correlated with learning noise.
Conclusions:
- Behavioral variability in reward-guided learning is primarily driven by limited computational precision, not active exploration.
- The locus coeruleus-norepinephrine system may modulate this learning noise.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Observational Learning
Instinctive Drift
Randomized Experiments
Simple randomization
Simple...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...

