Related Experiment Video
Updated: Apr 15, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement Learning
Abstract:
The increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase.
Related Concept Videos
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Robbers Cave
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Multi-input and Multi-variable systems
In the absence of...