Related Experiment Videos
Parametrized Graph Convolutional Multi-Agent Reinforcement Learning with Hybrid Action Spaces in Dynamic Topologies
Pei Chi1, Chen Liu2, Jiang Zhao2
1Institute of Unmanned System, Beihang University, Beijing 100191, China.
Abstract:
Multi-agent swarm collaboration, inspired by the collective behaviors of biological swarms in nature, has wide applications in dynamic open environments. However, hybrid action spaces in multi-agent reinforcement learning (MARL) present a critical challenge: the inherent coupling between discrete and continuous actions severely undermines policy stability and convergence, especially under dynamic topologies. Existing methods fail to decouple this coupling, leading to suboptimal policies and unstable training. This paper addresses the core problem of action coupling under dynamic topologies, proposing a Parametrized Graph Convolution Reinforcement Learning (P-DGN) method. Operating within the actor-critic framework, P-DGN decouples the optimization pathways for hybrid actions, with a biomimetic observation design inspired by starling flock behaviors: each agent only observes the states of its seven nearest neighbors to achieve efficient local interaction and global collaboration. Its actor network uses multi-head attention to build dynamic relation kernels, develops temporal relation regularization (TRR) to improve policy consistency across time steps, and generates continuous actions with a Gaussian policy. Meanwhile, P-DGN's critic network, based on deep Q-network (DQN), evaluates Q-values for discrete actions to guide optimal choices. We evaluate P-DGN in two different multi-agent cooperative environments. Experimental results show that compared with parametrized deep Q-network (P-DQN) and DQN baseline, the proposed method has faster convergence speed and stronger training stability. Moreover, with dense rewards, P-DGN agents learn emergent tactics like encirclement. Overall, P-DGN offers a new approach for optimizing hybrid action spaces in multi-agent systems within open, dynamic environments, balancing theoretical generality with practical utility, and its biomimetic design provides a biologically plausible framework for multi-agent swarm collaboration.
Related Concept Videos
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Multi-input and Multi-variable systems
In the absence of...
State Space Representation
Consider an RLC circuit, a...
Collisions in Multiple Dimensions: Introduction
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example: