在演员-关键算法中增强探索:一种鼓励可信的新状态的方法
IEEE transactions on cybernetics
|December 9, 2025
概括
演员-批判 (AC) 算法通过可信的新性来改善探索,这是探索具有高潜在学习益处的状态的内在奖励. 这提高了深度强化学习的样本效率和培训绩效.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度强化学习学习 (deep reinforcement learning) 是一种深度强化学习的方法.
背景情况:
- 演员-关键 (AC) 算法是有效的无模型深度强化学习方法.
- 有效地利用样本进行勘探和开采对于AC算法成功至关重要.
- 当前的方法往往无法量化新状态对政策学习的有用性,导致效率低下的探索.
研究的目的:
- 引入一种内在的奖励机制,称为可信的新性,以增强在AC算法中的探索.
- 通过激励探索具有高潜在学习效益的状态来提高样本效率和整体培训绩效.
主要方法:
- 开发了基于状态新性及其对政策学习的潜在实用性的内在奖励信号.
- 将可信的新性奖励集成到政策之外的关键演员算法中.
- 评估了对基准深度强化学习环境的拟议方法.
主要成果:
- 拟议的方法证明了样本效率的大幅提高.
- 在多个环境和算法中,训练回报率平均提高了19%.
- 标准偏差减少了30%,表明训练更加稳定.
结论:
- 可信的新性有效地增强了演员-关键算法的探索.
- 这种方法导致样本效率和培训绩效的显著提高.
- 这种方法为推进深度强化学习技术提供了一个有希望的方向.
相关概念视频
Optimal Foraging
13.5K
How animals obtain and eat their food is called foraging behavior. Foraging can include searching for plants and hunting for prey and depends on the species and environment.
13.5K
Actor-Observer Effect
303
The actor-observer effect, a cognitive bias closely linked to the fundamental attribution error, refers to the tendency for individuals to attribute their behavior to external, situational factors while explaining others’ behavior in terms of internal, dispositional traits. This asymmetry in attribution significantly influences social perception and judgment.Cognitive Mechanisms Behind the EffectTwo primary psychological mechanisms contribute to the actor-observer effect: differences in...
303
Observational Learning
795
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
795
Self-Evaluation: Self-Enhancement and Self-Verification
5.7K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.7K
Decision Making: P-value Method
6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K
Naturalistic Observations
16.9K
If you want to understand how behavior occurs, one of the best ways to gain information is to simply observe the behavior in its natural context. However, people might change their behavior in unexpected ways if they know they are being observed. How do researchers obtain accurate information when people tend to hide their natural behavior? As an example, imagine that your professor asks everyone in your class to raise their hand if they always wash their hands after using the restroom. Chances...
16.9K


