Related Experiment Video
Updated: May 27, 2025

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
8.7K
Hierarchical task network-enhanced multi-agent reinforcement learning: Toward efficient cooperative strategies.
Xuechen Mu1, Hankz Hankui Zhuo2, Chen Chen3
1School of Mathematics, Jilin University, Changchun, 130012, Jilin, China.
Summary
Hierarchical Symbolic Multi-Agent Reinforcement Learning (HS-MARL) enhances exploration in sparse reward environments. This novel approach significantly outperforms existing methods, particularly in challenging suboptimal settings.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Multi-agent reinforcement learning (MARL) faces significant challenges in environments with sparse rewards, often leading to premature exploration halts.
- Suboptimal exploration strategies hinder the performance of MARL agents in complex scenarios.
Purpose of the Study:
- To introduce Hierarchical Symbolic Multi-Agent Reinforcement Learning (HS-MARL), a novel framework designed to mitigate exploration difficulties in MARL.
- To reduce the effective exploration space by incorporating hierarchical knowledge and symbolic representations.
Main Methods:
- HS-MARL decomposes the state space using intermediate states and represents domain knowledge via the Hierarchical Domain Definition Language (HDDL) and the option framework.
- An enhanced Hierarchical Task Network (HTN) planner, pyHIPOP+, generates action sequences, which are assigned as policy functions by a high-level meta-controller.
- Intrinsic rewards are computed by the meta-controller to train symbolic option policies and refine the pyHIPOP+ heuristic function.
Main Results:
- HS-MARL demonstrated significant performance improvements over 15 state-of-the-art algorithms in sparse reward and suboptimal environments.
- Ablation studies confirmed the critical contributions of HS-MARL's intrinsic reward mechanism and the pyHIPOP+ component to its effectiveness.
- The approach was validated in both simulated sparse reward environments and a real-world football match scenario.
Conclusions:
- HS-MARL offers a robust solution for navigating complex MARL environments with sparse rewards and suboptimal conditions.
- The integration of hierarchical knowledge, symbolic options, and intrinsic rewards is key to enhancing exploration and agent performance.
- The developed framework provides a promising direction for advancing MARL research and applications.

