Related Experiment Video
Updated: Nov 27, 2025

Operation of the Collaborative Composite Manufacturing CCM System
Published on: October 1, 2019
The Convergence of a Cooperation Markov Decision Process System
Xiaoling Mo1, Daoyun Xu1, Zufeng Fu1,2
1College of Computer Science and Technology, Guizhou University, Guiyang 550025, China.
This study introduces a Cooperation Markov Decision Process (CMDP) for multi-agent learning. The research demonstrates that the value function converges, independent of initial conditions, and presents an algorithm for optimal strategy pairs.
Area of Science:
- Artificial Intelligence
- Reinforcement Learning
- Game Theory
Background:
- Traditional Markov Decision Processes focus on single agents.
- Multi-agent systems are increasingly prevalent in complex applications.
- Cooperative decision-making presents unique challenges in agent interactions.
Purpose of the Study:
- Introduce a Cooperation Markov Decision Process (CMDP) for two-agent systems.
- Analyze the convergence properties of the value function in CMDPs.
- Develop an algorithm to find optimal cooperative strategies.
Main Methods:
- Formulation of a two-agent Cooperation Markov Decision Process.
- Theoretical analysis of value function convergence.
- Development of an algorithm for optimal strategy pair identification.
Main Results:
- The value function in the CMDP system converges.
- Convergence value is independent of the initial value function.
- An algorithm successfully identifies optimal strategy pairs.
Conclusions:
- The CMDP framework effectively models cooperative decision-making between two agents.
- The convergence properties ensure stable learning outcomes.
- The proposed algorithm provides a method for achieving optimal cooperation.
More Related Videos
06:18The Collective Trust Game: An Online Group Adaptation of the Trust Game Based on the HoneyComb Paradigm
Published on: October 20, 2022
07:42An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Cooperative Allosteric Transitions
Cooperative Allosteric Transitions
Cooperative Allosteric Transitions
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....