Related Experiment Video
Updated: Jul 16, 2026

Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
Memory length and space shape multi-agent Q-learning dynamics
Wei Wang1, Xiaogang Li1, Yongjuan Ma1
1School of Statistics and Mathematics, Yunnan University of Finance and Economics, Kunming 650221, China.
Abstract:
In repeated interactions, players adjust their behavior based on previous moves. Higher memory leads to an exponential growth in the number of strategies, meaning players require more complex cognitive abilities. Even if players can observe the strategies and payoffs of their co-players, strategies imitated through social learning fail to guarantee effective responses to co-players to bring higher payoffs. We depict the human learning process through reinforcement learning, whereby players learn based on past experiences, independently of the strategies and payoffs of co-players. Here, we explore how different memory lengths and spaces, namely, memory-n, reactive-n, and reactive-n counting, affect the evolution of cooperation among reinforcement learning players. We found that memory-n players maintained higher cooperation than reactive-n players. Notably, higher memory promotes cooperation in memory-n players but inhibits it in reactive-n players. Reactive-n counting players can alleviate the negative effects of excessive memory by compressing the memory. Strategies with the nature of mutual cooperation and retaliation are key for reinforcement learning players to maintain cooperation. Our research highlights that judiciously adjusting the information available to players more effectively fosters cooperation within multi-agent systems.
Related Concept Videos
Collisions in Multiple Dimensions: Introduction
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Multi-input and Multi-variable systems
In the absence of...
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Associative Learning
Classical conditioning, also known...
Observational Learning