Related Experiment Videos
Efficient Multiagent Reinforcement Learning With OTOS: Online Tuning From an Offline Self-Organization Model
None:
The inefficiency of online learning poses a significant challenge in multiagent reinforcement learning (MARL). Offline-to-online learning can improve efficiency by learning from available offline data, which can provide important prior insights. However, this approach is often obstructed by the issue of incomplete offline data, which are a prevalent problem in practical scenarios. In these cases, offline data collection does not cover complete state or action trajectories for some or all agents. This article addresses the challenge of inefficient online strategy training with incomplete offline data by introducing a novel method called online policy tuning based on offline policy self-organization (OTOS). OTOS comprises two main modules: offline policy augmentation learning based on policy self-organization (OFS) and online policy automatic tuning (ONT). To address the limitations of incomplete datasets, OFS employs offline policy augmentation learning inspired by a famous self-organization model for swarm intelligence, i.e., the boid model. This approach encompasses three essential components: intention aggregation, policy alignment, and policy separation, corresponding to the three basic self-organization rules in the boid model. For the policy and evaluation networks trained offline, ONT leverages the actor-network learned offline to initialize the actor-network for online training. In addition, the offline critic network provides supplementary value signals to enhance the stability and efficiency of the online learning process. To further optimize learning, an automatic optimal temperature adjustment factor is designed, which dynamically adjusts the weight of these supplementary value signals. The proposed OTOS method has been extensively validated across four distinct football scenarios, consistently demonstrating superior performance.
Related Concept Videos
Observational Learning
Multi-input and Multi-variable systems
In the absence of...
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Automatic Processing and Automatic Social Behavior
Associative Learning
Classical conditioning, also known...