在密集采样的纵向评估中,游戏化强化学习的可靠性
Monja P Neuser1, Anne Kühnel1,2,3, Franziska Kräutlein1
1Department of Psychiatry and Psychotherapy, Tübingen Center for Mental Health, University of Tübingen, Tübingen, Germany.
PLOS digital health
|September 6, 2023
概括
这项研究介绍了Influenca,这是一款用于重复奖励学习评估的应用程序. 强化学习参数显示可靠性差至相当低,强调需要更好的个人学习模型.
科学领域:
- 认知神经科学 认知神经科学
- 计算精神病学是一种计算精神病学.
- 行为经济学是一种行为经济学.
背景情况:
- 强化学习对激励至关重要,在精神障碍中发生变化.
- 目前基于实验室的奖励学习评估限制了测量频率和测试复试可靠性.
- 准确地测量个体学习对于理解决策和心理健康至关重要.
研究的目的:
- 推出Influenca,这是一个开源应用程序,用于反复,生态有效的奖励学习评估.
- 用重复测量来评估强化学习参数和心理状态的可靠性.
- 为量化基于价值的决策中的个人内部和个人间差异提供一个游戏化的框架.
主要方法:
- 开发Influenca,这是一个跨平台应用程序,具有新的奖励学习任务.
- 整合生态瞬间评估 (EMA) 以实时跟踪心理和生理状态.
- 对384名参与者的强化学习参数 (学习率,奖励敏感性) 和心理状态项目的分析.
主要成果:
- 强化学习参数表现出较差或相当的类内相关性 (ICCs:0.22-0.53),表明在主体内和主体之间存在显著的差异.
- 心理状态评估项目显示ICC与强化学习参数相比.
- 在重复评估中观察到学习和决策参数的实质性变化.
结论:
- Influenca提供了一种游戏化和可定制的方法来优化奖励学习的重复评估.
- 这项研究强调了强化学习参数在学科内部和学科之间存在相当大的变化.
- 这些发现强调了需要改进的方法来可靠地量化基于价值的决策的个人差异随着时间的推移.
相关概念视频
Longitudinal Research
12.0K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.0K
Reinforcement Schedules
203
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
203
Reinforcement
274
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
274
Longitudinal Studies
187
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
187
Reliability and Validity
12.8K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.8K
Observational Learning
209
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
209


