Related Experiment Videos
A Fully Decentralized Online Actor-Critic Algorithm for Constrained Multiagent Reinforcement Learning Over Directed
Abstract:
This article studies constrained multiagent reinforcement learning (C-MARL) over fixed directed graphs, where agents cooperatively optimize a globally averaged long-term objective while satisfying globally averaged constraints. To address the challenges of decentralized learning with a row-stochastic communication matrix, we propose a fully decentralized online actor-critic (AC) algorithm with Perron-vector normalization. Each agent updates its policy parameters, critic parameters, and local multiplier estimates based on local information received from its in-neighbors. Convergence guarantees are established under linear function approximation for the critic. Numerical simulations demonstrate the effectiveness and robustness of the proposed algorithm.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Actor-Observer Effect
Observational Learning