在人类强化学习中,预测错误与学习率之间的非线性关系
Boluwatife Ikwunne1, Jolie Parham1, Erdem Pulcu1,2,3
1Psychopharmacology and Emotion Research Lab, Department of Psychiatry, University of Oxford, Oxford, United Kingdom.
PLoS computational biology
|September 12, 2025
概括
这项研究揭示了强化学习模型中预测错误和学习率之间的新型非线性关系. 一个指数对数函数解释了预测错误如何即时更新学习率,影响代理适应.
科学领域:
- 计算神经科学是一种神经科学.
- 机器学习 机器学习
- 认知科学 认知科学
背景情况:
- 强化学习 (RL) 模型解释了动态环境中的适应性行为.
- 在RL中,预测错误 (PEs) 和学习率之间的准确关系尚未得到充分理解.
研究的目的:
- 调查在RL中PEs和学习率之间的非线性关系.
- 引入一种新型的RL模型,其中包含了PEs和学习率的指数-逻辑函数.
主要方法:
- 模拟RL代理人的模拟.
- 重新分析现有的RL数据集.
- 一个新的实验来测试拟议的模型.
主要成果:
- 在模拟,数据集再分析和新实验中证明了PEs和学习率之间的非线性关系.
- 开发了一个指数对数函数来建模PE即时转化为学习率的模型.
- 在信念更新期间积累的学习速率的观察到的生理相关性,与模型预测保持一致.
结论:
- 这项研究引入了一种新的非线性模型,用于理解RL中的PE学习率动态.
- 这个框架提供了一个更准确的表现,即代理人如何适应不断变化的环境.
- 研究结果表明,基于预测错误的学习速度调整背后的生理机制.
相关概念视频
Timing and Consequences on Behavior
360
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
360
Reinforcement Schedules
462
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
462
Nonlinear Pharmacokinetics: Causes of Nonlinearity
711
Nonlinearity in drug pharmacokinetics is caused by various factors influencing how a drug is absorbed, distributed, metabolized, and excreted. Understanding these nonlinear processes is crucial for predicting drug behavior in the body and optimizing drug dosing regimens.
Nonlinear drug absorption can occur when the process is rate-limited by solubility, carrier-mediated transport systems, or saturation of the presystemic gut wall or hepatic metabolism. For instance, high doses of riboflavin...
Nonlinear drug absorption can occur when the process is rate-limited by solubility, carrier-mediated transport systems, or saturation of the presystemic gut wall or hepatic metabolism. For instance, high doses of riboflavin...
711
Reinforcement
842
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
842
Regression Toward the Mean
6.9K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.9K
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K


