強化学習における選択ヒステリーズの進化:ポジティビティバイアスの適応値と漸進的な忍耐力の比較
Isabelle Hoxha1,2, Léo Sperber1,2, Stefano Palminteri1,2
1Département d'Etudes Cognitives, École Normale Supérieure, Université de Recherche Paris Sciences et Lettres, Paris 75005, France.
まとめ
過去の選択を繰り返すような 補強学習のバイアスは 適応的かもしれません ポジティビティバイアスは 進化的に安定した環境で 段階的な選択の持続とは違います
科学分野:
- 認知科学
- 計算神経科学
- 行動経済学
背景:
- 強化学習の実験では 過去の選択を繰り返す傾向が よく見られます
- この選択の繰り返しはしばしば非対称的な更新や段階的な選択の継続によって説明される.
- メタアナリシスは人間の補強学習におけるこれらのメカニズムを確認したが,それらの適応的価値は比較されないままである.
研究 の 目的:
- 補強学習における選択の繰り返しの基礎となる計算プロセスの適応的価値を調査する.
- 不対称な更新の進化的安定性 (ポジティビティバイアス) と,さまざまな環境における漸進的な選択の持続性を比較する.
主な方法:
- 様々なシミュレーション環境で進化アルゴリズムを用いた 強化学習エージェントをシミュレートする.
- 異なる選択-繰り返しのメカニズムの進化的安定性と強さを評価した.
主要な成果:
- 多数のシミュレーションシナリオにおいて,非対称的な更新として現れるポジティブバイアスは進化的に安定していることが判明しました.
- 漸進的な選択の持続性の出現は,ポジティブ性バイアスと比べると,一貫性や強度が弱かった.
- 環境の文脈はこれらのバイアスの適応選択に大きく影響する.
結論:
- ポジティビティバイアスのような補強学習における計算バイアスは,進化的選択によって適応し,好まれる可能性があります.
- これらのバイアスの適応的価値は環境特異であり,行動と生態学的圧力の間の微妙な相互作用を強調しています.
- ポジティビティバイアスは テストされた環境における 漸進的な選択の持続性よりも 進化的強さを示しています
関連する概念動画
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Timing and Consequences on Behavior
152
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
152
Hindsight Biases
3.9K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.9K
Generalization, Discrimination, and Extinction
784
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
784
Reinforcement
341
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
341
Instinctive Drift
314
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
314


