奖励预测错误信号的动态编码在强化学习期间,在子腹部 tegmental 区域的奖励预测错误信号
Zhigang Shang1,2, Jiashuo Zhang1,2, Mengmeng Li1,2
1School of Electrical and Information Engineering, Zhengzhou University, Zhengzhou 450001, China.
eNeuro
|February 19, 2026
概括
子大脑在腹部 tegmental 区域 (VTA) 的活动显示在学习过程中奖励预测错误 (RPE) 信号. 神经信号随着学习的进展而转向预测线索,类似于哺乳动物.
科学领域:
- 神经科学是一个神经科学.
- 比较心理学比较心理学
- 行为神经科学 行为神经科学
背景情况:
- 奖励预测错误 (RPEs) 对于学习至关重要,腹膜区域 (VTA) 活动与哺乳动物的RPE信号相关.
- 在强化学习期间的鸟类VTA动态不太了解,需要进行跨物种比较.
研究的目的:
- 在强化学习任务期间调查子的VTA神经活动.
- 描述VTA动态,特别是类似RPE的信号,如何随着鸟类的学习而演变.
主要方法:
- 在使用微电线阵列的子中记录了VTA多单元活动 (MUA) 在提示指导操作任务期间.
- 在培训期间,分析的MUA与任务事件 (线索和结果) 保持一致.
主要成果:
- 随着学习稳定,VTA活动显示出与学习相关的转变:结果锁定调制减少,而提示锁定调制随着学习稳定而增加.
- 这种VTA活动的时间转移与时间差异学习模型和RPE信号一致.
结论:
- 子 VTA 聚合的 MUA 展示了与学习相关的动态,暗示了 RPE 处理.
- 这些发现支持在脊椎动物物种中保存基本的多巴胺学习机制.
相关概念视频
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Operant Conditioning
3.0K
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
3.0K
Reinforcement
978
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
978
Reinforcement Schedules
542
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
542
Avoidance Learning and Learned Helplessness
2.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Timing and Consequences on Behavior
484
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
484


