Learning to ground multi-agent reinforcement learning with masked multi-agent AutoEncoding
Mingxiao Feng1, Lin Liu1, Wengang Zhou1
1CAS Key Laboratory of GIPAS, University of Science and Technology of China, Hefei, China; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, China.
Abstract:
Data efficiency is a persistent bottleneck in multi-agent reinforcement learning (MARL), especially under partial observability where each agent must act on limited local observations. Existing approaches often struggle to fully exploit the shared latent structure across agents, leading to slow convergence. We propose Masked Multi-Agent Autoencoding (MMAAE), a novel self-supervised representation learning framework for centralized training with decentralized execution (CTDE) MARL. MMAAE introduces a dual-space masking mechanism that simultaneously performs random masking and reconstruction in both the raw observation space and the learned representation space. By employing an asymmetric autoencoder to reconstruct masked observation tokens and aligning reconstructed features with target representations via a contrastive objective, MMAAE forces the encoder to capture robust inter-agent dependencies and global state information. Crucially, this auxiliary task is optimized jointly with the policy, requiring no additional environment interactions. We integrate MMAAE with popular MARL backbones (MAPPO and MAT) and evaluate it across diverse benchmarks: SMAC, Multi-Agent MuJoCo, and MAQC. Empirical results demonstrate that MMAAE significantly improves sample efficiency, reducing Time-to-Threshold by up to 42% and increasing early-training AUC by up to 38%, while matching or exceeding the final performance of strong baselines. Ablation studies further confirm that the dual-masking strategy is essential, as removing either observation- or representation-space masking leads to a marked drop in performance.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Associative Learning
Classical conditioning, also known...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
