Hindsight-based state space exploration via counterfactual intrinsic reward assignment

Yuchen Liang1, Qinchen Yang2, Fukai Zhang2

  • 1College of Artificial Intelligence, Xi'an Jiaotong University, Xi'an, Shaanxi, 710049, China.

Summary

This study introduces a counterfactual intrinsic reward method for reinforcement learning. It uses past experiences to guide agents toward novel states, improving exploration efficiency.

Related Concept Videos

Hindsight Biases01:12

Hindsight Biases

Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
Counterfactual Thinking01:19

Counterfactual Thinking

Counterfactual thinking is a cognitive process wherein individuals mentally reconstruct alternative versions of past events, often beginning with “what if” or “if only.” This reflective mechanism plays a significant role in shaping emotional experiences and guiding future behavior. Though typically triggered by unfavorable or unexpected outcomes, counterfactual thinking can also emerge in mundane, everyday decisions and experiences, revealing its deep entrenchment in human cognition.Types of...
Decision Making: P-value Method01:09

Decision Making: P-value Method

The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Fundamental Attribution Error01:14

Fundamental Attribution Error

According to some social psychologists, people tend to overemphasize internal factors as explanations—or attributions—for the behavior of other people. They tend to assume that the behavior of another person is a trait of that person, and to underestimate the power of the situation on the behavior of others. They tend to fail to recognize when the behavior of another is due to situational variables, and thus to the person’s state. This erroneous assumption is called the fundamental attribution...
The Anchoring-and-Adjustment Heuristic01:25

The Anchoring-and-Adjustment Heuristic

In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the $2,000...