在强化学习中演变选择歇斯底里:比较积极性偏差和逐渐坚持的适应价值
Isabelle Hoxha1,2, Léo Sperber1,2, Stefano Palminteri1,2
1Département d'Etudes Cognitives, École Normale Supérieure, Université de Recherche Paris Sciences et Lettres, Paris 75005, France.
概括
强化学习中的偏见,如重复过去的选择,可能是适应性的. 积极性偏差在许多环境中是进化稳定的,与逐渐选择的坚持不同.
科学领域:
- 认知科学
- 计算神经科学
- 行为经济学
背景情况:
- 强化学习实验常常显示出过去选择的重复趋势,
- 这种选择的重复往往是由不对称的更新或逐步选择的坚持解释的.
- 一项元分析证实了人类强化学习中的这些机制,但它们的适应性价值仍然是不可比拟的.
研究的目的:
- 研究强化学习中选择重复的基础计算过程的适应性价值.
- 为了比较不对称更新的进化稳定性 (积极性偏差) 与不同环境的逐渐选择持久性.
主要方法:
- 在各种模拟环境中使用进化算法模拟的强化学习代理.
- 评估不同选择重复机制的进化稳定性和稳定性.
主要成果:
- 在许多模拟场景中,表现为不对称更新的积极性偏差被发现是进化稳定的.
- 与积极性偏差相比,逐渐选择坚持的出现不那么一致和强大.
- 环境环境对这些偏差的适应性选择有很大的影响.
结论:
- 强化学习中的计算偏差,如积极性偏差,可以通过进化选择进行适应和偏好.
- 这些偏差的适应性价值是环境特有的,突出显示了行为和生态压力之间的微妙相互作用.
- 积极性偏差表现出更大的进化强度,而不是在经过测试的环境中逐渐选择的坚持.
相关概念视频
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Timing and Consequences on Behavior
152
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
152
Hindsight Biases
3.9K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.9K
Generalization, Discrimination, and Extinction
784
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
784
Reinforcement
341
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
341
Instinctive Drift
314
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
314


