Related Experiment Video
Updated: Jun 29, 2026

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
Published on: August 25, 2023
Hindsight-based state space exploration via counterfactual intrinsic reward assignment
Yuchen Liang1, Qinchen Yang2, Fukai Zhang2
1College of Artificial Intelligence, Xi'an Jiaotong University, Xi'an, Shaanxi, 710049, China.
This study introduces a counterfactual intrinsic reward method for reinforcement learning. It uses past experiences to guide agents toward novel states, improving exploration efficiency.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Causal Inference
Background:
- State space exploration is crucial but challenging in reinforcement learning.
- Current intrinsic reward functions often lack environmental knowledge, leading to inaccurate rewards.
- Hindsight knowledge from agent experience can improve intrinsic reward design.
Purpose of the Study:
- To develop a novel counterfactual intrinsic reward assignment method.
- To leverage hindsight knowledge for more effective state space exploration.
- To enhance reinforcement learning agent performance by identifying high-novelty states.
Main Methods:
- Utilizes counterfactual reasoning from causal learning to mine hindsight knowledge.
- Samples counterfactual actions and predicts next states using a structural causal model.
- Assigns intrinsic rewards based on the novelty of the actual state versus counterfactual states.
Main Results:
- Demonstrates improved state space exploration capabilities.
- Shows performance enhancements across 10 OpenAI Gym environments.
- Effectively guides agents to identify and reach novel states.
Conclusions:
- The counterfactual intrinsic reward method significantly improves reinforcement learning exploration.
- Leveraging hindsight and causal inference offers a promising direction for intrinsic reward design.
- The proposed method provides a robust framework for discovering novel states in complex environments.
Related Concept Videos
Hindsight Biases
Counterfactual Thinking
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Observational Learning
Fundamental Attribution Error
The Anchoring-and-Adjustment Heuristic
