奖励预测错误加权的提示突出性在模拟中产生上行为,具有不对称的学习和更的延迟折扣
Shivam Kalhan1, Marta I Garrido2, Robert Hester1
1University of Melbourne, School of Psychological Sciences, Melbourne, Victoria, Australia.
概括
这项研究引入了一个新的突出性数学模型来解释成行为. 它整合了多巴胺在学习和动机中的作用,成功模拟了诸如渴望和冲动等关键成特征.
科学领域:
- 神经科学是一个神经科学.
- 计算精神病学是一种计算精神病学.
- 行为经济学是一种行为经济学.
背景情况:
- 成行为与学习和动机系统的功能障碍有关.
- 现有的多巴胺在成中的作用模型在解释关键行为特征方面存在局限性.
- 之前的模型由Redish (2004) 和Zhang等. (2009) 专注于价值预测错误和动机,但未能解释成的关键方面.
研究的目的:
- 提出一种新的突出性的数学定义,将多巴胺在学习和动机中的作用整合起来.
- 模拟成行为的主要特征,包括对负面后果的敏感性降低,冲动性,渴望和非药物依赖.
- 为了建立在概念理论上,突出度调节内部表示,在成中更新.
主要方法:
- 开发了一个新的数学模型,在强化学习框架内定义突出.
- 利用单个参数模式来模拟成行为.
- 将模拟结果与现有模型和主要成特征进行比较.
主要成果:
- 提出的突出模型成功模拟了以前模型也产生的行为 (张等人,2009年;Redish,2004年).
- 该模型展示了模拟负面预测错误减轻权重的能力,更的延迟折扣,渴望,以及非药物成的方面.
- 这些发现支持了突出性调节内部表示更新的理论,有助于成行为.
结论:
- 突出性的新数学模型为多巴胺在学习中的双重作用和成中的动机提供了统一的解释.
- 这种综合方法更好地解释了上的关键特征,这些特征在以前的模型中没有被捕捉到.
- 这项研究强调了多巴胺的学习和激励方面之间的相互作用,通过影响内部表征的突出机制.
相关概念视频
Timing and Consequences on Behavior
101
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
101
Drug Abuse and Addiction: Pharmacological Phenomena
489
Drug dependence, abuse, and addiction are complex phenomena that can precipitate various abnormal states. Physical dependence refers to a state of pharmacological adaptation to a drug. This adaptation often results in tolerance—a reduced response to the drug after repeated administrations. When the drug use is abruptly stopped, withdrawal symptoms occur due to the body's need to readjust from the pharmacologically induced imbalance. However, tolerance and withdrawal symptoms do not...
489
Real-World Application of Classical Conditioning
598
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
598
Instinctive Drift
229
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
229


