皮多巴胺适应了从行动中学习的速度
Luke T Coddington1, Sarah E Lindo2, Joshua T Dudman3
1Howard Hughes Medical Institute, Janelia Research Campus, Ashburn, VA, USA. coddingtonl@hhmi.org.
Nature
|January 18, 2023
概括
大脑中的多巴胺信号调节行为政策的直接学习, 不仅仅是奖励预测. 这一发现扩大了动物行为和学习的强化学习模型.
科学领域:
- 神经科学
- 计算神经科学
- 动物的行为
背景情况:
- 人工智能和机器人利用政策和价值学习. 梅索林比克多巴胺在动物中用于奖励预测,但其在直接政策学习中的作用不太清楚.
- 强化学习模型解释了动物的行为,但精确的神经机制,特别是多巴胺在直接政策学习中的作用,需要进一步阐明.
研究的目的:
- 研究小鼠学习过程中的行为政策.
- 确定中边缘多巴胺在直接政策学习与价值学习中的作用.
- 测试多巴胺调节政策学习速度的神经网络模型.
主要方法:
- 在小鼠中全面分析面部和身体的运动,学习痕迹调节范式.
- 对多巴胺反应的个体差异与行为政策的出现进行关联.
- 在生理学上校准的多巴胺和与神经网络模型的比较.
主要成果:
- 初始多巴胺反应的个体差异与学习的行为政策相关,而不是价值编码.
- 多巴胺操纵效应与价值学习不一致,但由多巴胺调整学习速度的模型预测.
- 已表明相性多巴胺活性可以调节行为政策的直接学习.
结论:
- 在调节行为政策的直接学习中,中游 dopamine 起着至关重要的作用.
- 这一发现挑战了多巴胺仅仅作为奖励预测错误信号的传统观点.
- 这项研究提供了证据支持多巴胺在强化学习中发挥更广泛的作用,特别是在适应性政策更新中.
相关概念视频
Cognitive Learning
473
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
473
Purposive Learning
180
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
180
01:20Learned Behavior II
Learned Behavior IITeaching a dog to sit or learning how to ride a bike are examples of learned behaviors—actions acquired through experience and practice. These behaviors are not instinctive; instead, they develop over time as individuals interact with their environment. While some behaviors are automatic (like blinking), others are learned over time by watching, practicing, or being trained.Animals learn from their parents, their environment, and sometimes from trial and error. Whether a...
01:19Learned Behavior I
Learned Behavior ILearned behaviors are actions that animals develop through experience, observation, or practice rather than being born with them. For example, a dog learning to roll over or a baby bird figuring out how to crack open a seed are both learned behaviors. Unlike instincts, learned behaviors aren’t something you're born knowing. You pick them up through life and experience.Animals, including you, learn in all sorts of ways, such as copying others, solving problems, or remembering...
Observational Learning
255
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
255
Timing and Consequences on Behavior
139
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
139


