Related Experiment Video
Updated: Oct 2, 2025

11:18
Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
10.5K
Dynamical systems as a level of cognitive analysis of multi-agent learning: Algorithmic foundations of
1School of Mathematics, University of Leeds, Leeds, UK.
Neural Computing & Applications
|February 28, 2022
Summary
This study links evolutionary game theory and reinforcement learning to understand multi-agent learning dynamics. It proposes a framework to unify cognitive analysis levels, offering new insights into collective learning and decision-making under uncertainty.
Area of Science:
- Cognitive Science
- Machine Learning
- Game Theory
Background:
- Multi-agent learning dynamics are often analyzed using a dynamical systems perspective, linking evolutionary game theory and reinforcement learning.
- Existing interpretations of dynamical systems in multi-agent learning lack clarity, necessitating a more structured approach.
Purpose of the Study:
- To embed dynamical systems descriptions of multi-agent learning within different cognitive analysis levels.
- To clarify the connections between these abstraction levels for enhanced insight into multi-agent learning.
- To demonstrate the framework's utility with temporal-difference reinforcement learning.
Main Methods:
- Utilizing a dynamical systems framework to analyze multi-agent learning.
- Connecting evolutionary game theory and reinforcement learning.
- Proposing an on-line sample-batch temporal-difference algorithm with memory-batch and separated state-action value estimation.
Main Results:
- The deterministic dynamical systems description of temporal-difference reinforcement learning adheres to a minimum free-energy principle.
- This approach unifies boundedly rational game theory with decision-making under uncertainty.
- The proposed algorithm provides a micro-foundation for deterministic learning equations, with learning trajectories converging to deterministic equations under large batch sizes.
Conclusions:
- Embedding dynamical systems within cognitive analysis levels offers a robust framework for understanding multi-agent learning.
- The framework guides the full utilization of the dynamical systems approach in multi-agent learning research.
- This work bridges theoretical insights from game theory and practical applications in reinforcement learning.
Related Concept Videos
Observational Learning
357
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
357
Classification of Systems-I
348
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
348
Cognitive Learning
692
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
692
Classification of Systems-II
253
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
253
Multi-input and Multi-variable systems
185
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
185
Linear time-invariant Systems
508
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
508

