多巴胺过渡物编码奖励预测错误,独立于学习率
Andrew Mah1, Carla E M Golden1, Christine M Constantinople1
1Center for Neural Science, New York University, New York, NY, USA.
Cell reports
|October 12, 2024
概括
核心核中的多巴胺编码奖励预测错误 (RPEs),但不是学习率,这表明多巴胺独立的机制驱动了大鼠的动态学习.
科学领域:
- 神经科学是一个神经科学.
- 计算精神病学是一种计算精神病学.
- 强化学习是一种强化学习.
背景情况:
- 强化学习理论提出多巴胺编码奖励预测错误 (RPEs) 通过学习率进行缩放.
- 据信,由多巴胺调节的皮质状突触可塑性可以更新价值表示.
- 这个框架意味着多巴胺释放反映了RPE和学习率的产物.
研究的目的:
- 为了研究多巴胺如何在波动的环境中编码核核 (NAcc) 中的学习速度.
- 确定多巴胺释放是否反映了RPE和动态学习率.
主要方法:
- 老鼠执行了一项具有半可观察状态和不同奖励的任务.
- 行为分析检查了试验启动速度作为RPE的函数.
- 计算建模评估了学习率和隐藏状态的贝叶斯推理.
- 在NAcc中释放的多巴胺在任务执行期间被测量.
主要成果:
- 鼠根据RPE调整了试验启动速度,反映了动态的学习速度.
- 在状态过渡后,学习率增加,并随着对隐藏状态的信念更新而扩展.
- 在NAcc编码的RPEs中的多巴胺释放独立于计算的学习率.
结论:
- 在NAcc中的多巴胺编码奖励预测错误,但不是学习率,在波动的环境中.
- 有证据表明,多巴胺独立的机制负责实例化动态学习速率.
- 这些发现挑战了多巴胺在扩大学习速度中的作用的传统观点.
相关概念视频
Instinctive Drift
193
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
193
Timing and Consequences on Behavior
83
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
83


