Related Experiment Videos
MAB-PGC: Multiagent Mutual Awareness Belief via Progressive Graph Construction
Abstract:
In multiagent reinforcement learning (MARL), agents engage in collaborative or competitive interactions, adapting their policies to maximize rewards. However, in environments with only team rewards, agents cannot accurately perceive the rewards obtained from their individual actions. To address this limitation, we propose a multiagent belief state computation method termed mutual awareness belief (MAB). MAB extends belief modeling to decentralized partially observable settings by incorporating agents' behavioral intentions and collaborative relationships. This allows each agent to infer the contribution of individual behaviors to team rewards during decision-making and supports team-advantageous joint decisions. In addition, to address the issue of information lag, we develop progressive graph construction (PGC), a novel module for computing MAB. PGC constructs a dynamic graph with progressively updated node features and cooperation-informed edge weights. Through interaction with PGC, agents iteratively update and retrieve graph information to derive the MAB. Empirical results show that our method achieves competitive performance against baseline approaches on the StarCraft unit micromanagement benchmark and proves effective in the nonmonotonic Pursuit environment.