Related Experiment Video
Updated: Aug 23, 2026

A Flexible Platform for Monitoring Cerebellum-Dependent Sensory Associative Learning
Published on: January 19, 2022
DECHRL: Empowerment-driven delay-aware causal hierarchical reinforcement learning
Chenran Zhao1, Dianxi Shi2, Haotian Wang1
1College of Computer Science and Technology, National University of Defense Technology, Changsha, China.
None:
Many real-world tasks involve delayed effects, where the outcomes of actions emerge after varying time lags. Existing delay-aware reinforcement learning methods often rely on state augmentation, prior knowledge of delay distributions, or access to non-delayed data-limiting their generalization. Hierarchical reinforcement learning, by contrast, inherently offers advantages in handling delays due to its hierarchical structure, yet existing methods are restricted to fixed delays. To address these limitations, we propose Delay-Empowered Causal Hierarchical Reinforcement Learning (DECHRL). DECHRL consists of two modules: causal delay distribution modeling and delay-aware empowerment-driven hierarchical policy training. The former learns Granger-causal relationships from delayed interactions, while the latter uses this information to construct and optimize hierarchical policies. By guiding exploration toward controllable and informative states, DECHRL improves exploration efficiency and decision stability under temporal uncertainty. We evaluate DECHRL on modified 2D-Minecraft and MiniGrid environments with constructed stochastic delays, and further extend the evaluation to the low-dimensional continuous-state PointMaze task. Experimental results show that DECHRL effectively models causal delay distributions and significantly outperforms baselines under temporal uncertainty. The additional PointMaze results further demonstrate the applicability of the proposed delay-aware causal modeling mechanism to low-dimensional continuous navigation tasks with stochastic delayed transitions.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Purposive Learning
Associative Learning
Classical conditioning, also known...
Observational Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example: