Related Experiment Videos
Dynamic-layer transformer-based reinforcement learning for observation-constrained multi-agent roundup scenarios
Xizhao Li1, Ning Xu1, Qingjia Chi2
1School of Information Engineering, Wuhan University of Technology, Wuhan, 430070, Hubei, China.
Abstract:
To address the challenge of cooperative roundup of maneuvering targets under limited perception, this paper proposes TransMARL, a transformer-based multi-agent reinforcement learning framework for observation-constrained coordination. The roundup task is formulated as a decentralized partially observable Markov decision process (Dec-POMDP), together with a local observation model and a dynamically updated interaction graph. The proposed framework combines a graph feature encoding module with a policy execution module to support decentralized decision-making under partial observability. A task-informed reward function is designed to encourage angular coverage, target approach, formation uniformity, and collision avoidance. In addition, the transformer depth is adaptively adjusted according to the team size as an empirically motivated design choice to balance representational capacity and computational cost. Experimental results in a 2D obstacle-free simulation environment show that, under the evaluated settings, TransMARL achieves competitive and often improved performance relative to the selected baselines, especially under constrained sensing radii. These results suggest that the proposed framework is a practical and scalable approach for cooperative control in observation-constrained multi-agent roundup scenarios, while its broader generalization and formal theoretical characterization remain to be further studied.
Related Concept Videos
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Reinforcement Schedules
Once a behavior is learned,...
Associative Learning
Classical conditioning, also known...