Related Experiment Videos
Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization
Anton V Dozhdikov1, Arseniy M Sitkovskiy1,2
1Federal Research Sociological Centre of the Russian Academy of Sciences, Moscow, Russia.
Abstract:
Demographic systems face unprecedented challenges from simultaneous crises. Conventional statistical demography techniques and agent-based models often struggle to capture nonlinear inter-regional interactions during periods of severe socio-economic disruption. To address this, we propose MADDPG-EVO-DGM, a hybrid algorithm that integrates multi-agent deep reinforcement learning with evolutionary optimisation and meta-learning principles to model regional demographic processes under multiple crisis scenarios. Each region is treated as an autonomous agent learning to steer demographic policy levers, while periodic evolutionary "boosters" overcome local optima via population-based perturbations of actor network parameters. Additionally, a Darwin-Gödel Machine-inspired meta-learning mechanism adapts the booster triggers, enabling self-improvement in the learning process. We evaluate MADDPG-EVO-DGM on a simulation environment calibrated with real demographic data for eight federal regions of the Russian Federation over the period 2000-2024 and subject to ten concurrent crisis scenarios (e.g., pandemic, geopolitical conflict, economic collapse). Experiments demonstrate significantly faster convergence and improved performance over a baseline MADDPG: the hybrid approach achieves a higher final average reward (252.57 vs. 243.07) and 3.4 × lower convergence variance (σ = 0.24 vs. 0.80), indicating more reliable training. It also exhibits qualitative performance jumps of +68% during evolutionary phases and maintains 35%-45% greater resilience under crisis shocks compared to the baseline. To our knowledge, this is the first application of multi-agent reinforcement learning to large-scale demographic modeling under crises, opening new possibilities for evidence-based, crisis-resilient population policy design. Code, data, and logs are provided to ensure reproducibility.
Related Concept Videos
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Multi-input and Multi-variable systems
In the absence of...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Lagrange Multipliers: Problem Solving
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Evolutionary Psychology