Related Experiment Video
Updated: Nov 23, 2025

A Conflict Model of Reward-seeking Behavior in Male Rats
Published on: February 20, 2019
Are methamphetamine users compulsive? Faulty reinforcement learning, not inflexibility, underlies decision making in
Alex H Robinson1, José C Perales2, Isabelle Volpe3,4,5
1Turner Institute for Brain and Mental Health, Monash University, Melbourne, Victoria, Australia.
Abstract:
Methamphetamine use disorder involves continued use of the drug despite negative consequences. Such 'compulsivity' can be measured by reversal learning tasks, which involve participants learning action-outcome task contingencies (acquisition-contingency) and then updating their behaviour when the contingencies change (reversal). Using these paradigms, animal models suggest that people with methamphetamine use disorder (PwMUD) may struggle to avoid repeating actions that were previously rewarded but are now punished (inflexibility). However, difficulties in learning task contingencies (reinforcement learning) may offer an alternative explanation, with meaningful treatment implications. We aimed to disentangle inflexibility and reinforcement learning deficits in 35 PwMUD and 32 controls with similar sociodemographic characteristics, using novel trial-by-trial analyses on a probabilistic reversal learning task. Inflexibility was defined as (a) weaker reversal phase performance, compared with the acquisition-contingency phases, and (b) persistence with the same choice despite repeated punishments. Conversely, reinforcement learning deficits were defined as (a) poor performance across both acquisition-contingency and reversal phases and (b) inconsistent postfeedback behaviour (i.e., switching after reward). Compared with controls, PwMUD exhibited weaker learning (odds ratio [OR] = 0.69, 95% confidence interval [CI] [0.63-0.77], p < .001), though no greater accuracy reduction during reversal. Furthermore, PwMUD were more likely to switch responses after one reward/punishment (OR = 0.83, 95% CI [0.77-0.89], p < .001; OR = 0.82, 95% CI [0.72-0.93], p = .002) but just as likely to switch after repeated punishments (OR = 1.03, 95% CI [0.73-1.45], p = .853). These results indicate that PwMUD's reversal learning deficits are driven by weaker reinforcement learning, not inflexibility.
Related Concept Videos
Drug Abuse and Addiction: Pharmacological Phenomena
Attention-Deficit/Hyperactivity Disorder
Diagnostic Criteria and Symptoms
To diagnose ADHD, symptoms must manifest before age 12 and be evident across multiple settings....
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Instinctive Drift
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...

