Related Experiment Videos
Deterministic dynamics of distributional multi-agent reinforcement learning
Clémence Bergerot1,2, Pawel Romanczuk1,3, Wolfram Barfuss4,5,6
1Department of Biology, Humboldt Universität zu Berlin, Berlin, Germany.
Abstract:
Understanding how cognition shapes behavior across contexts remains a fundamental challenge for many disciplines. In particular, for the optimism heuristic-i.e., the tendency to overweight positive (relative to negative) information-knowledge remains fragmented, with models developed in specific domains in isolation. Here, we present a unifying computational framework by deriving the deterministic dynamics of distributional multi-agent reinforcement learning. Our approach discretizes return distributions through a finite set of neurons, consistent with recent empirical findings on distributional coding in the brain. We validate our framework by reproducing established results across three iconic domains spanning individual bandit choice under resource variability, social coordination, and risky choice. Beyond validation, we uncover novel interactions among optimism, return discretization, and temporal discounting. Specifically, we identify conditions under which return discretization generates choice hysteresis and, in extreme parameter regimes, inescapable perseveration. We further reveal "individual dilemmas": circumstances where agents gravitate toward suboptimal yet stable strategies, offering a mechanistic explanation for incoherent choice patterns. Our framework bridges neuroscience, psychology, and collective behavior, enabling empirically testable hypotheses about how cognitive biases propagate from individual cognition to social outcomes in complex environments.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Distribution Reliability and Automation
Observational Learning
Multi-input and Multi-variable systems
In the absence of...
Dynamic Equilibrium