Related Experiment Video
Updated: Jun 20, 2026

Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
Decoding species coexistence: A reinforcement learning perspective
Kaiwen Jiang1, Chenyang Zhao1,2, Shengfeng Deng1
1Shaanxi Normal University, School of Physics and Information Technology, Xi'an 710061, People's Republic of China.
Abstract:
A central goal in ecology is to understand how biodiversity is maintained. Previous theoretical works have employed the rock-paper-scissors (RPS) game as a toy model, demonstrating that population mobility is crucial in determining the species' coexistence. One key prediction is that biodiversity is jeopardized and eventually lost when mobility exceeds a certain value-a conclusion at odds with empirical observations of highly mobile species coexisting in nature. To address this discrepancy, we introduce a joint reinforcement learning framework to study a spatial RPS model, where individuals' mobility for each species is not fixed but is guided by a common experience pool in the form of a Q-table, and its members jointly revise it via a Q-learning algorithm. Our results show that all three species can coexist stably, with extinction probabilities remaining low across a broad range of baseline migration rates. Mechanistic analysis reveals that individuals develop two behavioral tendencies: survival priority (escaping from predators) and predation priority (remaining near prey). While species coexistence emerges from the balance of the two tendencies, their imbalance jeopardizes biodiversity. Notably, there is a symmetry breaking of action preference in a particular state that is responsible for the divergent species densities. Furthermore, when Q-learning species interact with fixed-mobility counterparts, those with adaptive mobility exhibit a significant evolutionary advantage. Our study suggests that joint reinforcement learning offers a promising perspective for uncovering the mechanisms of biodiversity and designing conservation strategies.
Related Concept Videos
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Associative Learning
Classical conditioning, also known...
Understanding Species and Reproductive Barriers
Predator-Prey Interactions
