Related Experiment Video
Updated: Dec 17, 2025

08:42
Assessment of Social Cognition in Non-human Primates Using a Network of Computerized Automated Learning Device ALDM Test Systems
Published on: May 5, 2015
12.5K
Combined model-free and model-sensitive reinforcement learning in non-human primates
Bruno Miranda1,2,3, W M Nishantha Malalasekera1, Timothy E Behrens4,5
1Institute of Neurology, Department of Clinical and Movement Neurosciences, University College London, London, United Kingdom.
Plos Computational Biology
|June 23, 2020
Summary
This study shows rhesus monkeys use a combination of model-free (MF) and model-sensitive (MS) reinforcement learning (RL) strategies, with a strong preference for MS methods in decision-making tasks.
Area of Science:
- Neuroscience
- Cognitive Science
- Computational Neuroscience
Background:
- Reinforcement learning (RL) theory distinguishes between model-free (MF) and model-sensitive (MS) strategies.
- MF strategies rely on reward history, while MS strategies utilize task dynamics knowledge.
- Combining MF and MS strategies offers complementary computational strengths, but empirical evidence, especially in non-human primates, is limited.
Purpose of the Study:
- To investigate the interplay of MF and MS reinforcement learning strategies in rhesus monkeys.
- To discriminate between MF and MS strategy use in a novel two-stage decision task.
- To analyze how task structure and reward history influence choice behavior and response vigor.
Main Methods:
- Rhesus monkeys were trained on a two-stage decision task.
- Descriptive analysis of choice behavior and response vigor was performed.
- Trial-by-trial computational modeling was used to analyze decision-making strategies, including a novel combined RL model.
Main Results:
- Both task structure and reward history significantly influenced choice and response vigor.
- Behavior was best explained by a combination of MF and MS strategies, with a dominant and persistent model-sensitive influence.
- A new RL model incorporating specific credit assignment weighting was developed to account for observed behavior.
- Response vigor showed distinct MF and MS influences.
Conclusions:
- Rhesus monkeys employ a sophisticated combination of reinforcement learning strategies.
- Model-sensitive strategies play a dominant role in decision-making for this species.
- These findings offer novel insights into the neural and computational mechanisms of reinforcement learning in primates.
Related Concept Videos
Observational Learning
726
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
726
Cognitive Learning
911
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
911
Reinforcement Schedules
379
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
379
Instinctive Drift
542
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
542

