Related Experiment Video
Updated: Aug 12, 2025

Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
Published on: August 4, 2023
Realistic Actor-Critic: A framework for balance between value overestimation and underestimation.
Sicen Li1,2, Qinyun Tang1,2, Yiming Pang1,2
1College of Mechanical and Electrical Engineering, Harbin Engineering University, Harbin, China.
Realistic Actor-Critic (RAC) improves reinforcement learning by balancing exploration and exploitation. This method enhances sample efficiency and performance, particularly in complex environments.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Reinforcement Learning
Background:
- Value approximation bias in reinforcement learning leads to suboptimal policies and hinders effective exploration-exploitation decisions.
- Existing algorithms struggle to provide a clear understanding of value bias impact and methods for stable, efficient exploration.
Purpose of the Study:
- To clarify the impact of value bias on reinforcement learning performance.
- To develop a method for efficient exploration with stable updates.
- To enhance sample efficiency in reinforcement learning algorithms.
Main Methods:
- Designed a simple episodic tabular MDP to study value underestimation and overestimation in actor-critic methods.
- Proposed the Realistic Actor-Critic (RAC) framework using Universal Value Function Approximators (UVFA).
- RAC learns policies with different value confidence bounds within a single neural network to manage under/overestimation trade-offs.
Main Results:
- Identified that fixed hyperparameter settings can cause agents to over-explore low-value states.
- RAC enables directed exploration using upper bounds and avoids overestimation with lower bounds.
- Achieved 10x sample efficiency and 25% performance improvement over Soft Actor-Critic in the Humanoid environment.
Conclusions:
- Provides insights into the exploration-exploitation trade-off by analyzing policy access to low-value states under varying confidence bounds.
- Proposes RAC as a unified framework to enhance sample efficiency in continuous control domains when combined with current actor-critic methods.
More Related Videos
05:21Characterization of the Sense of Agency over the Actions of Neural-machine Interface-operated Prostheses
Published on: January 7, 2019
13:04Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods
Published on: September 19, 2012
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
Lazarus's Cognitive Appraisal Theory
Primary Appraisal:...
Self-Discrepancy Theory
Self-Presentation: Self-Monitoring and Self-Handicapping
Fundamental Attribution Error
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...