多巴胺作用预测错误作为一个无价值的教学信号
Francesca Greenstreet1, Hernando Martinez Vergara1,2, Yvonne Johansson1
1Sainsbury Wellcome Centre for Neural Circuits and Behaviour, University College London, London, UK.
Nature
|May 14, 2025
概括
小鼠通过两种多巴胺信号学习:一种是奖励,另一种是重复行为. 行动预测错误强化了条纹体中的重复学习, 通过奖励信号来建立稳定的关联.
科学领域:
- 神经科学
- 动物的行为
- 计算神经科学
背景情况:
- 动物的选择行为包括寻求奖励和重复行为.
- 多巴胺信号,包括奖励预测错误和动作预测错误,被认为是为了强化这些不同的行为倾向.
研究的目的:
- 研究与运动相关的多巴胺活动在条形体中的作用.
- 确定行动预测错误是否作为学习的无价值教学信号.
主要方法:
- 在小鼠的听觉区分任务.
- 测量与运动相关的斑纹体尾部的多巴胺活性.
- 导致神经回路的操纵.
- 计算模型.
主要成果:
- 排列体尾部的运动相关多巴胺活动编码了行动预测错误.
- 行动预测错误作为一个无价值的教学信号,加强重复的关联.
- 单独的动作预测错误不支持奖励导向学习,但当与奖励预测错误电路配对时会巩固关联.
结论:
- 两种不同的多巴胺预测错误 (奖励和动作) 协同作用以支持学习.
- 这些信号在不同的状区域加强了不同类型的关联,有助于灵活而稳定的行为选择.
相关概念视频
Time-Domain Interpretation of PD Control
74
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
74
Purposive Learning
93
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
93
Drugs Affecting Neurotransmitter Synthesis
1.2K
Drugs affecting neurotransmitter synthesis can impact the adrenergic neuron and the synthesis of neurotransmitters. For example, α-methyltyrosine and carbidopa target specific enzymes involved in catecholamine synthesis. α-methyltyrosine inhibits the enzyme tyrosine hydroxylase, which converts tyrosine into dopamine. By blocking this enzyme, α-methyltyrosine reduces dopamine production and other catecholamines. Carbidopa, on the other hand, inhibits the enzyme dopa decarboxylase,...
1.2K
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Law of Effect
1.3K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.3K


