相关实验视频
Updated: Jun 6, 2025

06:57
Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
10.9K
人类强化学习中缓慢变化的特征的诱导偏见
Noa L Hedrich1,2,3, Eric Schulz4,5, Sam Hall-McMaster1,6
1Max Planck Research Group NeuroCode, Max Planck Institute for Human Development, Berlin, Germany.
PLoS computational biology
|November 25, 2024
概括
当重要的线索慢慢变化时,人类会更好地学习. 这项研究发现,人体强化学习的偏见有利于奖励预测缓慢变化的特征,提高学习效率.
科学领域:
- 认知科学 认知科学
- 神经科学是一个神经科学.
- 机器学习 机器学习
背景情况:
- 有效的行为需要在新环境中识别与目标相关的特征.
- 预先了解奖励预测特征的知识,如它们的变化率,可以指导这个过程.
- 行为相关的过程往往变化比随机噪音慢.
研究的目的:
- 为了调查人类是否表现出一种偏见,当任务相关的特征变化缓慢而不是快速时,更有效地学习.
- 确定人类是否在强化学习中利用有关特征动态的先验知识.
主要方法:
- 295名人类参与者完成了两个实验和一个复制任务,包括从二维强盗那里学习奖励.
- 参与者学会了与盗的缓慢或快速变化的特征相关的奖励.
- 使用卡尔曼过器的计算建模分析了基于特征相关性和速度的学习率和调整.
主要成果:
- 当相关特征变化缓慢,不相关特征变化迅速时,参与者获得了更多的奖励.
- 在条件之间对未见的特征值的概括中没有发现显著差异.
- 人类学习率对于缓慢的特征更高,参与者根据特征相关性和速度调整学习.
结论:
- 人类强化学习表现出一种偏见,有利于奖励预测的更慢变化的特征.
- 这种偏见表明人类如何在动态环境中学习的具体策略.
- 学习速度的调整对环境特征的相关性和时间动态都很敏感.
相关概念视频
Instinctive Drift
187
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
187
Purposive Learning
99
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
99
Reinforcement Schedules
132
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
132
Timing and Consequences on Behavior
79
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
79
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Generalization, Discrimination, and Extinction
445
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
445

