Related Experiment Video
Updated: Mar 6, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
A Multi-Agent Continual reinforcement learning framework with multi-Timescale replay and dynamic task classification
Yang Liu1, Xiang Feng1, Huiqun Yu1
1Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, 200237, China.
None:
This paper proposes an innovative Multi-Agent Continual Reinforcement Learning (MACRL) framework to address the challenges of continual learning in dynamic multi-agent systems. Traditional reinforcement learning suffers from catastrophic forgetting and inefficient cross-task knowledge transfer in non-stationary environments. To overcome these limitations, we introduce two key components: (1) a Multi-Timescale Replay (MTR) buffer, which hierarchically stores experiences across varying timescales to balance new task learning and prior knowledge retention, and (2) a dynamic task classification mechanism that employs an attention-based contextual encoder to measure task similarity and adaptively route policies, thereby minimizing inter-task interference. Experiments on cooperative multi-agent benchmarks (LBF and PP) demonstrate that our framework achieves up to higher average return compared to baselines in sequential task learning, with superior zero-shot generalization performance. Ablation studies further validate the critical roles of MTR and task classification in mitigating catastrophic forgetting. This work provides a scalable solution for collaborative decision-making in complex, evolving environments.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Multi-input and Multi-variable systems
In the absence of...
Associative Learning
Classical conditioning, also known...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
