通过强化学习出现类似信念的表征
Jay A Hennig1,2, Sandra A Romero Pinto2,3,4, Takahiro Yamaguchi3,5
1Department of Psychology, Harvard University, Cambridge, Massachusetts, United States of America.
PLoS computational biology
|September 11, 2023
概括
动物可以使用循环神经网络 (RNN) 从不完整的信息中学习奖励价值. 这些网络在没有明确计算信念的情况下产生准确的奖励预测,为适应性行为提供可扩展的解决方案.
科学领域:
- 计算神经科学是一种计算神经科学.
- 动物行为 动物行为
- 机器学习是机器学习.
背景情况:
- 动物必须预测适应性行为的未来奖励 (价值).
- 强化学习模型是常见的,但动物往往学习不完整的状态信息.
- 之前的模型提出了对部分可观测任务的明确信念估计.
研究的目的:
- 为了调查反复神经网络 (RNN) 是否可以直接从观察中学习价值估计.
- 为了确定RNN是否可以在没有明确的信念计算的情况下产生生物学上可信的奖励预测错误.
- 探索RNN能力与信念信息编码之间的关系.
主要方法:
- 在具有部分可观测性的任务上训练一个循环神经网络 (RNN).
- 使用统计,功能和动态系统视角分析RNN的内部表示.
- 将RNN产生的奖励预测错误与实验观测进行比较.
主要成果:
- 在没有明确的信念计算的情况下,RNN学会了估计价值并产生奖励预测错误.
- 当网络容量足够大时,RNN的学习表示编码了信念信息.
- 这表明了在具有有限计算资源的系统中价值估计的潜在机制.
结论:
- 循环神经网络可以直接从观察中学习价值估计,模仿动物的学习.
- 在部分可观测的环境中,明确的信念计算并不总是必要的.
- 这种方法为理解复杂系统中的适应性行为提供了一个计算可扩展的替代方案.
相关概念视频
Observational Learning
207
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
207
Reinforcement
273
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
273
Cognitive Learning
307
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
307
Purposive Learning
139
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
139
Associative Learning
434
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
434
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K


