Related Experiment Video
Updated: Jun 29, 2026

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
Published on: August 25, 2023
Hindsight-based state space exploration via counterfactual intrinsic reward assignment
Yuchen Liang1, Qinchen Yang2, Fukai Zhang2
1College of Artificial Intelligence, Xi'an Jiaotong University, Xi'an, Shaanxi, 710049, China.
Abstract:
State space exploration has been a core challenge in reinforcement learning, due to the difficulty in designing an intrinsic reward function that can guide agents to precisely find high-novelty states in the state space. Most existing methods design the reward function without fully utilizing any knowledge concerning environments, which often produces inaccurate reward signals. Instead, this paper assumes the hindsight knowledge extracted from the agent's previous exploration experience can help intrinsic reward designs. Therefore, the counterfactual intrinsic reward assignment method is presented, which uses the counterfactual reasoning mechanism from the causal learning field to mine hindsight knowledge from previously explored states, and utilizes this knowledge for better intrinsic reward assignment to the agent. The core idea is to "backtrack" the agent's obtained exploration result and "reflect on" whether a more novel state could have been discovered, if the agent had selected another action at the time. Concretely, the method first samples a batch of "counterfactual" actions that differ from the current action from a policy, then uses a structural causal model to predict their corresponding next states. Subsequently, the method will give a large reward if the state actually explored is more novel than that in counterfactual reasoning, otherwise a low reward, thereby "rewarding" agents to effectively identify and find novel states in the state space. Simulations show the effectiveness in improving state space exploration, with the performance enhanced over 10 OpenAI Gym environments.
Related Concept Videos
Hindsight Biases
Counterfactual Thinking
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Observational Learning
Fundamental Attribution Error
The Anchoring-and-Adjustment Heuristic
