人类强化学习中的价值规范化的功能形式
Sophie Bavard1,2,3, Stefano Palminteri1,2
1Laboratoire de Neurosciences Cognitives et Computationnelles, Institut National de la Santé et Recherche Médicale, Paris, France.
eLife
|July 10, 2023
概括
奖励值取决于上下文,范围正常化,而不是分裂性正常化,解释了学习和决策中的这一现象. 这一发现为认知过程提供了新的见解.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 计算神经科学是一种神经科学.
背景情况:
- 在强化学习中,奖励处理是取决于背景的.
- 这种上下文依赖常常被划分性规范化解释,但范围规范化是一种替代机制.
- 之前的研究缺乏设计来区分这些规范化模型.
研究的目的:
- 为了调查分化或范围规范化是否更好地解释了上下文依赖的奖励表示.
- 设计一种能够区分这两种规范化理论的新型学习任务.
主要方法:
- 开发了一项新的学习任务,通过操纵跨语境的选项数量和价值范围来处理.
- 收集和分析行为数据.
- 使用计算建模来测试规范化账户.
主要成果:
- 行为和计算分析伪造了分裂的正常化账户.
- 证据强烈支持奖励表示的范围正常化规则.
- 该研究成功地区分了这两种规范化机制.
结论:
- 范围规范化,而不是分裂性规范化,是学习中上下文依赖的奖励表示的基础.
- 这些发现推动了我们对控制学习和决策上下文依赖的计算机制的理解.
相关概念视频
Reinforcement
280
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
280
Primary and Secondary Reinforcers
304
In psychology, reinforcement is a key concept in behavior modification. B.F. Skinner demonstrated this with his experiments involving rats in what is known as a Skinner box. The rats learned to press a lever to receive food, a primary reinforcer that fulfilled their innate need for nourishment.
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
304
Reinforcement Schedules
205
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
205
Generalization, Discrimination, and Extinction
632
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
632
Observational Learning
213
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
213
Real-World Application of Classical Conditioning
630
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
630


