Related Experiment Video
Updated: Sep 13, 2025

Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
Published on: June 1, 2015
Data-driven equation discovery reveals nonlinear reinforcement learning in humans
Kyle J LaFollette1,2, Janni Yuval3, Roey Schurr4
1Department of Psychological Sciences, Case Western Reserve University, Cleveland, OH 44106.
A new Quadratic Q-Weighted model improves reinforcement learning (RL) predictions by incorporating nonlinear dynamics and negativity biases. This advanced computational model offers better insights into human learning and decision-making compared to traditional linear approaches.
Area of Science:
- Cognitive Science
- Computational Neuroscience
- Behavioral Economics
Background:
- Reinforcement learning (RL) models are crucial for understanding human decision-making.
- Traditional RL models often use linear updates for reward expectations, which may oversimplify complex human behavior.
- There is a need for more nuanced computational models to capture the intricacies of learning and reward processing.
Purpose of the Study:
- To develop and validate a novel computational model of reinforcement learning (RL) that addresses limitations of traditional linear models.
- To explore the application of equation discovery algorithms in uncovering new RL models from behavioral data.
- To investigate the role of nonlinear dynamics and negativity biases in human reward prediction errors.
Main Methods:
- Utilized equation discovery algorithms, a method adapted from physics and biology, to identify potential RL models.
- Proposed a new model, the Quadratic Q-Weighted model, based on differential equations capturing linear and nonlinear functions.
- Tested the model's generalizability and predictive accuracy against classical RL models using nine published datasets.
Main Results:
- The Quadratic Q-Weighted model demonstrated superior predictive accuracy over traditional models in eight out of nine published datasets.
- The model revealed that reward prediction errors follow nonlinear dynamics and exhibit negativity biases.
- Findings indicate an underweighting of rewards when expectations are low and an overweighting of reward absence when expectations are high.
Conclusions:
- The Quadratic Q-Weighted model offers a more accurate and interpretable approach to modeling human learning and decision-making.
- This study highlights the power of integrating behavioral tasks with advanced computational methods for discovering cognitive patterns.
- The findings represent a significant advancement in developing broadly applicable and insightful computational models of human cognition.
Related Concept Videos
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Once a behavior is learned,...
Purposive Learning
Associative Learning
Classical conditioning, also known...
Observational Learning

