子中的动态学习率偏差:来自强化学习和神经相关的洞察力
Fuli Jin1,2, Lifang Yang1,2, Long Yang1,2
1School of Electrical and Information Engineering, Zhengzhou University, Zhengzhou 450001, China.
Animals : an open access journal from MDPI
|February 10, 2024
概括
子表现出动态的学习策略,在概率任务中将他们的学习率偏差从负向正转变. 这种行为变化与条纹体中的神经活动相关,为动物的强化学习提供了洞察力.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 动物行为 动物行为
背景情况:
- 强化学习研究表明,动物以不同的方式处理奖励预测错误.
- 在人类和动物中观察到学习速度偏差,但其在学习过程中的动态变化仍然不清楚.
研究的目的:
- 在子的概率学习任务中调查学习速度偏差的动态变化.
- 探索行为学习策略与鸟类状体中神经活动之间的关系.
主要方法:
- 在概率学习任务期间,在子条形体中记录了行为选择和局部场势 (LFPs).
- 应用了带有和没有学习率偏差的强化学习模型,以适应行为数据和估计选项值.
- 分析了条纹性LFP功率与模型估计的选项值之间的相关性.
主要成果:
- 在整个学习过程中,子的学习率偏差从负面转变为积极.
- 状马功率 (31-80 Hz) 与受动学习速率偏差影响的选项值相关.
- 行为和神经数据支持子的动态学习策略.
结论:
- 子使用动态学习策略,随着时间的推移调整他们的学习率偏差.
- 状神经活动,特别是马功率,反映了这种动态的学习过程.
- 这些发现为非人类动物的强化学习机制提供了宝贵的见解.
相关概念视频
Instinctive Drift
221
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
221
Cognitive Learning
243
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
243
Timing and Consequences on Behavior
94
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
94


