Related Experiment Video
Updated: Apr 16, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Multi-agent contrastive exploration via value decomposition discrepancy
Siying Wang1, Hongfei Du2, Chiyu Cai3
1School of Automation Engineering, University of Electronic Science and Technology of China, Chengdu, China; College of Computer Science and Artificial Intelligence, Southwest Minzu University, Chengdu, China; Intelligent Perception and Control Key Laboratory of Sichuan Province, Sichuan University of Science & Engineering, Zigong, China.
None:
Value decomposition, as a factorization approach in multi-agent reinforcement learning (MARL), has been influential in the development of many effective value-based algorithms. Existing studies on value decomposition that focus on representation capability often suffer from low sample efficiency, while other factorization approaches may also have limited representation capability, which may still prevent collaborators from discovering the optimal joint action. To enhance agents' exploration with an unlimited joint state-value function, we propose a Multi-Agent Contrastive Exploration method (MACE) leveraging value decomposition discrepancies and contrastive principles. MACE determines update weights based on the discrepancy between the different value decomposition estimates to set update weights and introduces this difference as an intrinsic target in the update process. Additionally, MACE designs an exploration preference network inspired by this difference, explicitly adjusting the exploration preferences of agents during interactions. Through the experiments of various types and difficulties on Matrix Games and Starcraft Multi-Agent Challenge, we show that MACE not only significantly outperforms the baselines in learning speed and final performance, but also effectively maintains the higher expressivity, representing an innovative solution that integrates the advantages of existing algorithms.
More Related Videos
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Related Concept Videos
Self-Discrepancy and Its Effects
Self-Discrepancy Theory
Causes of Similarity-Dissimilarity Effect
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...