Related Experiment Video
Updated: Sep 14, 2025

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
Representation-driven sampling and adaptive policy resetting for improving multi-Agent reinforcement learning
Weiqiang Jin1, Xingwu Tian2, Ningwei Wang1
1School of Information and Communications Engineering, Xi'an Jiaotong University, Xi'an, Shanxi, 710049, Shanxi, China.
We introduce eXJTU-MARL, a novel approach to multi-agent reinforcement learning (MARL) that enhances exploration and learning efficiency. This method overcomes common MARL challenges like policy convergence and sampling inefficiency in complex environments.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Multi-agent reinforcement learning (MARL) is crucial for complex decision-making but faces challenges like policy convergence and sampling inefficiency.
- Large state, observation, and action spaces exacerbate these issues, leading to suboptimal strategies and extensive training requirements.
Purpose of the Study:
- To propose a novel MARL approach, eXJTU-MARL, enhancing exploration and trajectory learning efficiency.
- To address insufficient exploration and sampling inefficiency in mainstream MARL algorithms.
Main Methods:
- Introduction of adaptive policy resetting to prevent premature convergence.
- Development of state representation-based balanced experience sampling for improved data efficiency.
- Implementation within the eXJTU-MARL framework for multi-agent decision-making tasks.
Main Results:
- eXJTU-MARL significantly enhances sample efficiency and exploration capabilities.
- The approach prevents premature convergence to suboptimal policies.
- Experiments in StarCraft Multi-Agent Challenge show superior performance over existing MARL baselines.
Conclusions:
- eXJTU-MARL effectively improves exploration and learning efficiency in complex MARL environments.
- Adaptive policy resetting and balanced experience sampling are key contributions.
- The proposed method offers a robust solution for advanced multi-agent decision-making.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Random Sampling Method
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Role of Shaping in Operant Conditioning
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
Observational Learning

