强化学习参数的测试-重新测试可靠性
Jessica V Schaaf1,2,3, Laura Weidinger4,5, Lucas Molleman6,5
1Department of Psychology, University of Amsterdam, Amsterdam, the Netherlands. jessica.schaaf@radboudumc.nl.
Behavior research methods
|September 8, 2023
概括
计算表型化使用模型参数来研究个体差异. 然而,这项研究发现强化学习模型参数的测试复试可靠性较差,这表明参与者变化,就像情绪一样,影响结果.
科学领域:
- 计算精神病学是一种计算精神病学.
- 认知神经科学是一种认知神经科学.
- 心理测量研究的研究.
背景情况:
- 计算表型利用计算模型参数来理解认知过程中的个体差异.
- 计算表型的一个关键假设是行为和模型参数的时间稳定性.
- 计算模型的测试-重新测试可靠性,特别是强化学习模型,在很大程度上没有特征.
研究的目的:
- 调查法定强化学习模型的测试-重新测试可靠性.
- 在两个常见的学习范式中评估模型参数的可靠性:双臂强盗任务和反转学习任务.
- 确定计算表型化是否是评估个体差异的可靠方法.
主要方法:
- 两个独立的队列 (N=69和N=47) 完成了在线学习任务.
- 在测试和重新测试会议之间使用了五周的时间间隔.
- 用模型参数的类内相关系数 (ICC) 评估了测试重复测试的可靠性,并与人格和认知指标进行了比较.
主要成果:
- 对于人格和认知指标 (ICCs: .67.93) 观察到高的测试复试可靠性.
- 强化学习模型参数的测试复试可靠性普遍较差 (盗任务ICCs: .02.52;反转学习任务ICCs: .01.71).
- 模拟证实了该研究能够检测出高可靠性的能力,表明糟糕的结果反映了真正的不稳定性.
- 发现参与者的情绪 (压力和快乐) 解释了参与者内部的一些变化.
结论:
- 由于参数稳定性差,这些发现挑战了计算精神病学当前计算表型化实践的可靠性.
- 参与者内部显著的变化,可能受到情绪等因素的影响,使模型参数的解释变得复杂.
- 计算表型的未来发展必须考虑和解决个体的变化,以确保强大和有意义的见解.
相关概念视频
Reinforcement Schedules
202
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
202
Reinforcement
273
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
273


