Related Experiment Videos
LLM-Augmented Multi-Agent Reinforcement Learning for Cross-Scenario Knowledge Transfer
Chao Li1, Yanfei Liu1, Jieling Wang1
1Department of Basic Courses, Rocket Force University of Engineering, Xi'an 710025, China.
None:
Multi-agent reinforcement learning (MARL) relies on trial-and-error interactions to update policies. However, trial-and-error learning typically requires extensive interactions to achieve satisfactory performance, resulting in low sample efficiency, which limits its application in the real world. To reduce the trial-and-error costs of MARL and accelerate the convergence of multi-agent collaborative policies, we propose a MARL policy transfer method named LoLM-MARL, based on fine-tuning large language models (LLMs). First, leveraging the general world knowledge and reasoning capabilities of LLMs, low-rank adaptation (LoRA) is employed to fine-tune the pre-trained model on source tasks, thereby providing general decision-making knowledge for cross-scenario policy transfer. Second, a dynamic prompt construction method for LLMs is designed. By dynamically eliminating the state information of ineffective agents from the prompts, the method provides denser observation data for the large language model, thereby enhancing its policy representation capability in specific complex collaborative scenarios. Meanwhile, the dynamic prompt design concept enriches the training sub-scenarios for the algorithm, thereby laying the foundation for the model to learn more general decision-making knowledge. Finally, a Kullback-Leibler (KL) divergence regularization method based on an annealing strategy is constructed to ensure consistency between the policy distributions of the fine-tuned model and the pre-trained model, effectively mitigating the catastrophic forgetting problem during the fine-tuning process of the pre-trained model. Experimental results show that in zero-shot transfer tasks, LoLM-MARL achieves a maximum improvement of 101.4% in average win rate compared to existing state-of-the-art (SOTA) methods. In six few-shot transfer tasks, our method consistently achieves better generalization performance than traditional SOTA methods, and improves the convergence speed by 4 to 30 times compared to the training-from-scratch approach, providing a new solution paradigm for efficient policy transfer in complex dynamic environments.
Related Concept Videos
Associative Learning
Classical conditioning, also known...
Observational Learning
Multi-input and Multi-variable systems
In the absence of...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example: