Related Experiment Video
Updated: Jun 28, 2025

07:43
Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
Published on: August 4, 2023
1.9K
Boosting On-Policy Actor-Critic With Shallow Updates in Critic
IEEE Transactions on Neural Networks and Learning Systems
|April 15, 2024
Summary
Least-squares deep policy gradient (LSDPG) combines batch and deep reinforcement learning for improved data efficiency. This hybrid approach enhances learning stability and sample efficiency in deep reinforcement learning tasks.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Deep Reinforcement Learning
Background:
- Deep reinforcement learning (DRL) utilizes deep neural networks (NNs) for function approximation.
- Batch reinforcement learning (BRL) offers stable training and data efficiency with fixed representations.
Purpose of the Study:
- To propose Least-Squares Deep Policy Gradient (LSDPG), a hybrid method combining BRL and DRL.
- To achieve enhanced stability and data efficiency by integrating least-squares methods with online DRL.
Main Methods:
- LSDPG employs a shared network for feature sharing between policy (actor) and value function (critic).
- It utilizes regularized least-squares temporal difference (LSTD) for policy evaluation in a stationary critic setting.
- An auxiliary task distills critic features into the representation for improved learning.
Main Results:
- The critic converges to a regularized TD fixpoint, and the actor converges to a locally optimal policy under specific conditions.
- LSDPG demonstrates improved sample efficiency compared to Proximal Policy Optimization and Phasic Policy Gradient on the Procgen benchmark.
Conclusions:
- LSDPG effectively combines the strengths of BRL and DRL, offering a stable and data-efficient approach.
- The method shows promise for improving performance in complex reinforcement learning environments.
Related Concept Videos
Observational Learning
168
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
168
Self-Evaluation: Self-Enhancement and Self-Verification
5.2K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.2K
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
Effects of feedback
550
Feedback in control systems plays a critical role in shaping various operational parameters, extending beyond simple error reduction to influence stability, bandwidth, gain, impedance, and sensitivity. Understanding these effects requires examining a basic feedback system characterized by defined input, output, error, and feedback signals.
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
550
Self-Discrepancy Theory
18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.3K
Social Facilitation
32.0K
Not all intergroup interactions lead to negative outcomes. Sometimes, being in a group situation can improve performance. Social facilitation occurs when an individual performs better when an audience is watching than when the individual performs the behavior alone. This typically occurs when people are performing a task for which they are skilled.
32.0K

