学习表达奖励预测 类似错误的多巴胺活性 需要时间的塑料表示
Ian Cone1,2,3, Claudia Clopath1, Harel Z Shouval2,4
1Department of Bioengineering, Imperial College London, London, United Kingdom.
Research square
|October 4, 2023
概括
时间差异学习面临着固定的时间假设的挑战. 一个新的FLEX模型提供灵活学习的错误,以更好地进行强化学习,与实验数据保持一致.
科学领域:
- 计算神经科学是一种计算神经科学.
- 机器学习 机器学习
- 神经生物学 神经生物学 神经生物学
背景情况:
- 时间差异 (TD) 强化学习是理解基于大脑的学习的主要框架.
- TD理论认为神经元代表奖励预测错误 (RPEs),即预期和实际奖励之间的差异.
- 腹部体区域的多巴胺基神经元表现出与TD模型中的RPE信号一致的发射模式.
结论:
- FLEX框架为强化学习提供了一种更具适应性和实验一致性的方法.
- 拟议的FLEX模型成功地解释了各种强化学习范式.
- 来自FLEX模型的结果与神经奖励处理的现有和新分析的实验数据保持一致并解释.
相关概念视频
Long-term Potentiation
55.3K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre- and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
55.3K
Timing and Consequences on Behavior
114
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
114
Purposive Learning
135
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
135
Associative Learning
428
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
428
Observational Learning
202
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
202


