Related Experiment Video
Updated: Jun 6, 2025

The HoneyComb Paradigm for Research on Collective Human Behavior
Published on: January 19, 2019
Navigating the unknown: Leveraging self-information and diversity in partially observable environments
Devdhar Patel1, Hava T Siegelmann1
1Manning College of Information and Computer Science, University of Massachusetts, Amherst, MA, 01003, USA.
Abstract:
Reinforcement learning algorithms often struggle to learn in partially observable environments, where different states of the environment may appear identical. However, not all partially observable environments pose the same level of difficulty for learning. This work introduces the concept of dissonance distance, a metric that can estimate the difficulty of learning in such environments. We demonstrate that self-information, such as internal oscillations or memory of previous actions, can increase the dissonance distance and make learning easier in partially observable environments. Additionally, sensory occlusion may occur after learning was completed, leading to a lack of sufficient information and catastrophic failure. To address this, we propose a spatially layered architecture (SLA) inspired by the brain, which trains multiple policies in parallel for the same task. SLA can change the amount of external information processed at each timestep, providing an adaptive approach to handle the changing information in the environment state-space. We evaluate the effectiveness of our SLA method showing learnability and robustness against realistic noise and occlusion in sensory inputs for the partially observable Continuous Mountain Car environment. We hypothesize that multi-policy approaches like SLA might explain the complex dopamine dynamics in the brain that cannot be explained with the state of the art scalar Temporal Difference error.
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
Naturalistic Observations
Self-Presentation: Self-Monitoring and Self-Handicapping
Observational Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Stereotype Threat and Self-fulfilling Prophecies

