快速:人们可以在竞争性游戏中以适应的方式利用无模型和基于模型的强化学习
1Kansas State University, Manhattan, KS, USA.
概括
人类通过使用适应性策略,在"石头,纸,剪刀"中表现优于强化学习 (RL) 的对手. 参与者改变了基于对手可预测性的方法,偏离了标准的RL预测.
科学领域:
- 行为经济学是一种行为经济学.
- 计算神经科学是一种计算神经科学.
- 游戏理论的游戏理论.
背景情况:
- 竞争性社会互动驱动战略学习.
- 有限的研究存在于对手可预测性如何影响人类的战略和性能随着时间的推移.
研究的目的:
- 根据对手的可预测性,研究人类的战略行为和绩效如何变化.
- 为了比较人类策略与无模型和基于模型的强化学习 (RL) 算法在岩石,纸,剪刀.
主要方法:
- 使用石头,纸张,剪刀进行了两项实验,人类参与者与RL编程的计算机对手进行了比赛.
- 实验1采用了无模型的RL算法,只更新了所选的动作值.
- 实验2使用基于模型的RL算法,更新选定的和未选定的动作值.
主要成果:
- 参与者在无模型和基于模型的RL对手中始终表现出色.
- 人类参与者采用了不同的策略:在实验1中赢-停留/输-转移,在实验2中赢-转移/输-转移.
- 观察到的人类策略与标准RL模型和学习理论的预测有所不同.
结论:
- 在竞争环境中,人类的战略决策比传统的RL模型预测的更具动态性和适应性.
- 强化学习作为理解战略决策的框架,但人类的适应性需要进一步研究.
- 未来的研究应该探索额外的计算模型和竞争游戏,以了解适应性战略转变.
相关概念视频
Avoidance Learning and Learned Helplessness
3.3K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
3.3K
Reinforcement
1.1K
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
1.1K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
387
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
387
Reinforcement Schedules
671
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
671
Randomized Experiments
9.3K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
9.3K
