Q値と環境認知に基づく強化学習における適応的探索戦略
Tenglong Yang1, Jingbao Hou2, Peiyi Zhang3
1National-local Joint Engineering Laboratory of Marine Mineral Resources Exploration Equipment and Safety Technology, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China; College of Mechanical and Electrical Engineering, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China.
まとめ
この研究は,探査と搾取をバランスにする強化学習のための適応的探査戦略であるVarを紹介しています. Varは学習のパフォーマンスを向上させ,様々な環境における災害的な行動を軽減します.
科学分野:
- 人工知能 (AI) とは,人工知能 (AI) のことです.
- 機械学習 (Machine Learning) とは,機械学習 (Machine Learning) について学ぶことです.
- 強化学習による学習です.
背景:
- 強化学習 (RL) は,連続的な意思決定に優れているが,探査と採掘のバランスに苦戦している.
- 既存の方法は,しばしば単一の信号を使用し,非効率的な探査と非最適のソリューションにつながります.
研究 の 目的:
- 強化学習のための適応的探索戦略を開発する.
- 現在の探査技術の限界に対処することによって,学習の効率とパフォーマンスを向上させる.
主な方法:
- 提案するVarは,Q値と環境認知 (Q値差と新奇性ボーナス) を利用した適応的探査戦略である.
- VarをQ学習 (Var-QL) とディープQネットワーク (Var-DQN) に統合する.
- LRU-Var-QLを大型タブラーディスクリート環境で導入し,最近使用された最小値 (LRU) キャッシュを組み込む.
主要な成果:
- Varは,最初は広範で深い探査を奨励し,搾取に移行します.
- LRU-Var-QLは,FrozenLake-v1.1.の破滅的なアクションが減少したことを実証しました.
- Var-QLとVar-DQNは,ベースラインと比較して,Atariゲームでより高い累積リターンとより速い学習を達成しました.
結論:
- 提案されたVar戦略は,強化学習の探索を強化します.
- Varベースの方法は,さまざまな環境における学習速度とパフォーマンスの有意な改善を示しています.
関連する概念動画
Cognitive Learning
1.4K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.4K
Decision Making: P-value Method
7.0K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.0K
Avoidance Learning and Learned Helplessness
2.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Environmental Influences on Intelligence
1.0K
Despite the strong genetic influence on traits like intelligence, environmental factors significantly shape outcomes. For example, while over 90% of height variation is due to genetic differences, environmental factors such as nutrition also have a notable impact. Similarly, for intelligence, changes in a child's surroundings can significantly alter their IQ. Research shows that enriched environments boost children's academic success and help them develop key cognitive skills. Children...
1.0K
Reinforcement
992
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
992


