政策纠正和国家条件的行动评估为少数人命终身深度强化学习
IEEE transactions on neural networks and learning systems
|April 30, 2024
概括
这项研究通过改进短暂的概括来增强终身深度强化学习 (DRL). 新方法显著提高了适应速度和对新任务的政策绩效.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度强化学习学习 (deep reinforcement learning) 是一种深度强化学习的方法.
背景情况:
- 终身深度强化学习 (DRL) 能够在没有知识损失的情况下不断适应新任务.
- 现有的终身DRL方法在不同任务的适应力和低于最佳的政策方面扎 (少数射击概括挑战).
研究的目的:
- 提出一种通用方法,以提高终身DRL方法,并具有少数射击概括能力.
- 解决当前终身DRL在适应新和显著不同任务方面的局限性.
主要方法:
- 选择性经验重复使用以改善适应培训.
- 在目标Q值上放松软max功能,以提高政策准确性.
- 测量和减少数据分布差异,以提高适应效率.
主要成果:
- 拟议的方法显著提高了六种代表性终身DRL方法的性能.
- 实验结果显示,通过这些方法,回报至少有25%的改善.
- 该方法提高了培训速度和政策的最佳性,而不是SOTA的DRL方法.
结论:
- 开发的通用方法有效地配备了现有的终身DRL方法,并进行了短暂的概括.
- 这项工作为终身DRL提供了显著的进步,克服了适应和泛化方面的关键挑战.
- 提出的技术可以在动态环境中提高学习效率和有效性.
更多相关视频
05:41A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
9.4K
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
7.5K
相关概念视频
Law of Effect
1.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.4K
Behavior Modification
142
Behavioral approaches have often been criticized for ignoring mental processes and focusing solely on observable behavior. However, these approaches provide an optimistic perspective for individuals seeking to change their behaviors. Rather than concentrating on intrinsic personality traits, behavioral approaches suggest that even longstanding habits can be modified by changing the reward contingencies that maintain them.
A real-world application of operant conditioning principles is applied...
A real-world application of operant conditioning principles is applied...
142
Role of Shaping in Operant Conditioning
302
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
302
Generalization, Discrimination, and Extinction
540
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
540
Reinforcement Schedules
144
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
144
Behaviorism
2.3K
The field of behaviorism was pioneered by figures such as Ivan Pavlov, John B. Watson, and B.F. Skinner fundamentally shifted the focus of psychology to the observable and controllable aspects of human and animal behavior. This shift marked a critical evolution in the discipline, emphasizing scientific rigor and experimental methodology.
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
2.3K
