メソリンビック・ドーパミンは 行動による学習の速度を調整します
Luke T Coddington1, Sarah E Lindo2, Joshua T Dudman3
1Howard Hughes Medical Institute, Janelia Research Campus, Ashburn, VA, USA. coddingtonl@hhmi.org.
Nature
|January 18, 2023
まとめ
ドーパミンの信号は 報酬の予測だけでなく 行動に関する直接的な学習を 制御します この発見は 動物の行動と学習の強化学習モデルを 拡張しています
科学分野:
- 神経科学
- 計算神経科学
- 動物 の 行動
背景:
- 人工知能とロボット工学は 政策と価値の学習を活用します メソリンビックドーパミンは動物における報酬予測のために研究されているが,直接的な政策学習におけるその役割はあまり理解されていない.
- 強化学習モデルは動物の行動を説明しますが,特にドーパミンの直接的な政策学習における役割は,さらなる解明を必要とします.
研究 の 目的:
- ネズミの学習過程で 行動方針がどのように進化するかを調べる
- 直接的な政策学習と価値学習におけるメソリンビック・ドーパミンの役割を決定する.
- ドーパミンが政策学習の速度を調節する ニューラルネットワークモデルをテストする
主な方法:
- トレースコンディショニングのパラダイムを学習するマウスの顔と体の動きの包括的な分析.
- ドーパミン反応の個々の差異と行動政策の出現を相関させる.
- メソリンビックドーパミンの生理学的に調整された操作とニューラルネットワークモデルとの比較.
主要な成果:
- 最初のドーパミンの反応の個別の違いは 学習された行動方針と相関していますが 値のエンコーディングではありません
- ドーパミンの操作効果は 価値学習と矛盾していますが ドーパミンが学習速度を調整するモデルによって予測されています
- 段階的なドーパミンの活動は,行動政策の直接的な学習を調節することが示されました.
結論:
- メソリンビック・ドーパミンは 行動政策の直接的な学習を 制御する上で重要な役割を果たします
- この発見は ドーパミンが報酬予測の 誤差信号にすぎないという 伝統的な見解に異議を唱えます
- この研究は,ドーパミンが強化学習,特に適応的政策の更新におけるより広範な役割を支持する証拠を提供します.
さらに関連する動画
関連する概念動画
Cognitive Learning
473
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
473
Purposive Learning
180
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
180
01:20Learned Behavior II
Learned Behavior IITeaching a dog to sit or learning how to ride a bike are examples of learned behaviors—actions acquired through experience and practice. These behaviors are not instinctive; instead, they develop over time as individuals interact with their environment. While some behaviors are automatic (like blinking), others are learned over time by watching, practicing, or being trained.Animals learn from their parents, their environment, and sometimes from trial and error. Whether a...
01:19Learned Behavior I
Learned Behavior ILearned behaviors are actions that animals develop through experience, observation, or practice rather than being born with them. For example, a dog learning to roll over or a baby bird figuring out how to crack open a seed are both learned behaviors. Unlike instincts, learned behaviors aren’t something you're born knowing. You pick them up through life and experience.Animals, including you, learn in all sorts of ways, such as copying others, solving problems, or remembering...
Observational Learning
255
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
255
Timing and Consequences on Behavior
139
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
139


