在人类背部流中基于奖励的选项竞争,以及从随机探索过渡到在连续空间中进行开发
Michael N Hallquist1, Kai Hwang2, Beatriz Luna3
1Department of Psychology, University of North Carolina, Chapel Hill, NC, USA.
Science advances
|February 23, 2024
概括
本研究提出了一个统一的强化学习模型,用于灵长类动物的感觉运动探索和利用. 该模型解释了动态背部流地图如何快速更新,并根据奖励反有效地从勘探到开发过渡.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 计算神经科学是一种神经科学.
背景情况:
- 灵长类动物使用动态的背部流图来执行感觉运动任务.
- 现有的模型 (强化学习,工作记忆) 为奖励编码和探索/利用提供了补充但不完整的解释.
- 需要一个统一的模型来解释快速地图可塑性和高效的决策.
研究的目的:
- 为了统一强化学习和工作记忆缓冲器模型的动态传感运动地图.
- 解释这些地图如何编码奖励,并解决勘探/开采困境.
- 调查地图更新和选项选择背后的神经机制.
主要方法:
- 开发一种新的强化学习模型,整合增量奖励学习和工作记忆.
- 使用神经成像 (可能是fMRI或EEG) 分析人类前双背流活动.
- 在决策任务中检查后部β/α振荡.
主要成果:
- 统一模型预测了快速,信息压缩地图更新和高效的勘探到开发的过渡.
- 前端平行脊流活动与竞争选项的数量相关,优先选项保持,替代品压缩.
- 后面的β/α振荡在发现有价值的新选项后迅速失同,这表明子群体编码.
结论:
- 统一模型成功地将现有的动态地图函数理论集成到背部流中.
- 脊流中的神经活动和振荡支持基于奖励的快速地图更新和偏见的选择选择.
- 这一框架促进了对灵长类动物的决策,学习和感觉运动控制的理解.
相关概念视频
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Timing and Consequences on Behavior
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Instinctive Drift
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...


