通过在关键方面进行浅薄的更新来促进政策上的行为者-批评者
IEEE transactions on neural networks and learning systems
|April 15, 2024
概括
最小方形深度政策梯度 (LSDPG) 结合了批量学习和深度强化学习,以提高数据效率. 这种混合方法提高了深度强化学习任务中的学习稳定性和样本效率.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度强化学习学习 (deep reinforcement learning) 是一种深度强化学习的方法.
背景情况:
- 深度强化学习 (DRL) 使用深度神经网络 (NN) 进行函数近似.
- 批强化学习 (BRL) 提供稳定的训练和数据效率与固定表示.
研究的目的:
- 提出最小平方深度政策梯度 (LSDPG),一种混合方法,将BRL和DRL结合起来.
- 通过将最小平方方法与在线DRL集成来实现更高的稳定性和数据效率.
主要方法:
- LSDPG采用一个共享网络,在政策 (行为体) 和价值功能 (关键) 之间进行功能共享.
- 它使用规范最小平方时间差 (LSTD) 在静止的批评环境中进行政策评估.
- 一个辅助任务将批评特征提炼到表达中,以改善学习.
主要成果:
- 批评者汇聚到一个规范化的TD固定点,演员在特定条件下汇聚到一个局部最佳的政策.
- 与Proximal Policy Optimization和Phasic Policy Gradient相比,LSDPG在Procgen基准上显示了更好的样本效率.
结论:
- LSDPG有效地结合了BRL和DRL的优势,提供了稳定和数据效率高的方法.
- 该方法在复杂的强化学习学习环境中有望提高性能.
相关概念视频
Observational Learning
168
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
168
Self-Evaluation: Self-Enhancement and Self-Verification
5.2K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.2K
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
Effects of feedback
550
Feedback in control systems plays a critical role in shaping various operational parameters, extending beyond simple error reduction to influence stability, bandwidth, gain, impedance, and sensitivity. Understanding these effects requires examining a basic feedback system characterized by defined input, output, error, and feedback signals.
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
550
Self-Discrepancy Theory
18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.3K
Social Facilitation
32.0K
Not all intergroup interactions lead to negative outcomes. Sometimes, being in a group situation can improve performance. Social facilitation occurs when an individual performs better when an audience is watching than when the individual performs the behavior alone. This typically occurs when people are performing a task for which they are skilled.
32.0K


