Related Experiment Video
Updated: Jun 6, 2025

The HoneyComb Paradigm for Research on Collective Human Behavior
Published on: January 19, 2019
Navigating the unknown: Leveraging self-information and diversity in partially observable environments.
Devdhar Patel1, Hava T Siegelmann1
1Manning College of Information and Computer Science, University of Massachusetts, Amherst, MA, 01003, USA.
This study introduces dissonance distance to measure reinforcement learning difficulty in partially observable environments. A novel spatially layered architecture (SLA) enhances learning robustness against sensory challenges.
Area of Science:
- Artificial Intelligence
- Computational Neuroscience
Background:
- Reinforcement learning (RL) algorithms face challenges in partially observable environments where states can be ambiguous.
- The difficulty of learning in such environments varies, but metrics to quantify this are lacking.
Purpose of the Study:
- Introduce 'dissonance distance' as a metric to estimate learning difficulty in partially observable environments.
- Propose a Spatially Layered Architecture (SLA) inspired by the brain to improve RL robustness against sensory input changes and occlusion.
Main Methods:
- Developed the concept of dissonance distance to quantify environmental learning difficulty.
- Proposed and implemented a Spatially Layered Architecture (SLA) training multiple policies in parallel.
- Evaluated SLA on the partially observable Continuous Mountain Car environment with realistic noise and occlusion.
Main Results:
- Demonstrated that self-information (e.g., internal oscillations, memory) increases dissonance distance, simplifying learning.
- SLA showed learnability and robustness against sensory noise and occlusion.
- SLA adaptively adjusts external information processing for changing environments.
Conclusions:
- Dissonance distance offers a novel way to understand and predict RL learning difficulty.
- The brain-inspired SLA provides an effective, robust solution for RL in dynamic, partially observable environments.
- Multi-policy approaches like SLA may offer insights into complex neural mechanisms, such as dopamine dynamics.
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
Naturalistic Observations
Self-Presentation: Self-Monitoring and Self-Handicapping
Observational Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Stereotype Threat and Self-fulfilling Prophecies

