数据驱动的方程发现揭示了人类的非线性强化学习
Kyle J LaFollette1,2, Janni Yuval3, Roey Schurr4
1Department of Psychological Sciences, Case Western Reserve University, Cleveland, OH 44106.
概括
一个新的方程Q权重模型通过结合非线性动态和消极偏差来改进强化学习 (RL) 预测. 与传统的线性方法相比,这种先进的计算模型为人类学习和决策提供了更好的洞察力.
科学领域:
- 认知科学 认知科学
- 计算神经科学是一种神经科学.
- 行为经济学是一种行为经济学.
背景情况:
- 强化学习 (RL) 模型对于理解人类决策至关重要.
- 传统的RL模型通常使用线性更新来预期奖励,这可能过于简化了复杂的人类行为.
- 需要更细微的计算模型来捕捉学习和奖励处理的复杂性.
研究的目的:
- 开发和验证一种新的强化学习 (RL) 计算模型,解决传统线性模型的局限性.
- 探索方程发现算法的应用,从行为数据中发现新的RL模型.
- 调查非线性动态和消极偏见在人类奖励预测错误中的作用.
主要方法:
- 利用方程发现算法,一种从物理和生物学中改编的方法,以识别潜在的RL模型.
- 提出了一种新的模型,即二进制Q加权模型,基于捕捉线性和非线性函数的微分方程.
- 使用九个已发表的数据集对经典RL模型进行了模型的概括性和预测准确性的测试.
主要成果:
- 在9个已发表的数据集中,8个中,方程Q权重模型表现出比传统模型更高的预测准确度.
- 该模型显示,奖励预测错误遵循非线性动态,并表现出消极偏见.
- 研究结果表明,当期望较低时,奖励的权重较低,当期望较高时,奖励缺席的权重过高.
结论:
- 方程Q权重模型提供了一个更准确,更易于解释的方法来建模人类的学习和决策.
- 这项研究强调了将行为任务与先进的计算方法整合在一起,以发现认知模式的力量.
- 这些发现代表了在开发人类认知的广泛适用和有洞察力的计算模型方面取得的重大进展.
相关概念视频
Law of Effect
1.6K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.6K
Reinforcement
343
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
343
Reinforcement Schedules
242
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
242
Purposive Learning
207
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
207
Associative Learning
579
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
579
Observational Learning
314
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
314


