在随机学习过程中延迟结果的信用分配的神经机制
Phillip P Witkowski1,2,3, Lindsay Rondot1,2, Zeb Kurth-Nelson4,5
1Center for Mind and Brain, University of California Davis, Davis, CA, 95618.
bioRxiv : the preprint server for biology
|August 16, 2024
概括
这项研究揭示了大脑如何为行为赋予信誉,即使经过长时间的延迟和干预决策. 侧轨前皮质和海马体是将选择与结果联系起来的关键,其辅助者是侧面前极皮质.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 决策 决策 决策 决策 决策
背景情况:
- 适应性行为需要将行动与结果联系起来.
- 对于特定的信用分配的神经机制仍然不清楚,特别是延迟和临时选择.
研究的目的:
- 调查特定信用分配的神经基础.
- 了解大脑如何将延迟的选择与复杂决策中的结果联系起来.
主要方法:
- 对fMRI数据的多变量模式分析.
- 一个适应性学习任务,涉及延迟的选择结果关联.
主要成果:
- 侧向轨道前皮层 (lOFC) 和海马体 (HC) 编码因果选择身份,以延迟信用分配.
- 侧面前极皮质 (lFPC) 支持在临时决策发生时的学习.
- lFPC保持悬而未决的选择表示,预测 lOFC/HC在信用分配期间的忠诚度.
结论:
- lOFC和HC对于学习选择结果关系与间接的延迟和决策至关重要.
- lFPC在延迟准确的信用分配期间保持选择信息方面发挥着至关重要的作用.
- 在LOFC/HC中及时恢复特定原因对于现实世界的学习和决策至关重要.
相关概念视频
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Timing and Consequences on Behavior
87
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
87
Cognitive Learning
233
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
233
Purposive Learning
106
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
106
Real-World Application of Classical Conditioning
540
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
540
Operant Conditioning
1.6K
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
1.6K


