背面-垂直强化学习网络连接和激励驱动的探索变化.
Ethan M Campbell1,2, Wanting Zhong3,4, Jeremy Hogeveen1,2
1Department of Psychology, University of New Mexico, Albuquerque, New Mexico 87131 ecampbell@unm.edu jhogeveen@unm.edu.
概括
更高的奖励减少了探索,并在决策中有利于基于模型的学习. 个人的认知特征和大脑连接性影响这种奖励驱动的平衡,影响探索-利用策略.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 决策科学 决策科学 决策科学
背景情况:
- 概率增强学习 (RL) 涉及在不确定性下做出决定,受内部模型 (基于模型) 或直接经验 (无模型) 的影响.
- 在平衡基于模型和无模型控制中的个体差异受到激励动机的影响,但可变奖励的影响未得到充分研究.
- 在RL控制策略上调节奖励效应的神经和认知因素在很大程度上是未知的.
研究的目的:
- 调查变量奖励激励如何影响决策中的基于模型和无模型学习之间的仲裁.
- 识别认知特征和神经特征的个体差异,以缓解奖励对RL控制的影响.
- 探索功能性大脑连接和奖励调节的探索-利用权衡之间的关系.
主要方法:
- 采用了两阶段的决策任务,使用不同的奖励激励措施.
- 使用计算建模,神经心理测试和神经影像 (静止状态功能连接) 进行了测试.
- 对两个独立的数据集进行了分析,包括两性.
主要成果:
- 增加奖励前景减少了探索,并将平衡转向基于模型的学习,在数据集中一致.
- 处理速度和分析思维风格调节了奖励对基于模型/无模型控制的影响.
- 在高激励下减少勘探,与RL电路的腹部 (刺激估值) 和背部 (动作估值) 之间的跨网络合增加相关.
结论:
- 奖励前景显著改变了勘探和开采之间的平衡,有利于基于模型的战略.
- 在RL网络中,认知特征和休息状态功能连接是奖励对决策影响的关键调节者.
- 在不断变化的激励下,腹部和背部RL网络之间的连接的完整性与适应性探索-开发调整有关.
相关概念视频
Instinctive Drift
175
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
175
Neuroplasticity
267
Neuroplasticity reflects the brain's remarkable capacity to adapt and evolve, responding dynamically to learning, experiences, or injury by reorganizing its neural circuitry. This reorganization involves creating new neural connections and refining old ones through a series of biological processes that contribute to the brain's lifelong development and adaptability.
267


