Related Experiment Video
Updated: May 30, 2026

Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
Spiking neural networks with different reinforcement learning (RL) schemes in a multiagent setting
Chris Christodoulou1, Aristodemos Cleanthous
1Department of Computer Science, University of Cyprus 75 Kallipoleos Avenue, P.O. Box 20537, 1678 Nicosia, Cyprus. cchrist@cs.ucy.ac.cy
Spiking agents trained with reinforcement learning (RL) achieved mutual cooperation in the Iterated Prisoner's Dilemma (IPD) game. Reward-modulated STDP enabled rapid learning and increased rewards, demonstrating effective multiagent learning.
Area of Science:
- Computational neuroscience
- Artificial intelligence
- Game theory
Background:
- Multiagent systems require sophisticated learning mechanisms for cooperation.
- Spiking neural networks (SNNs) offer biologically plausible models for intelligence.
- The Iterated Prisoner's Dilemma (IPD) is a standard benchmark for studying cooperation and competition.
Purpose of the Study:
- To investigate the efficacy of spiking agents trained with reinforcement learning (RL) in a multiagent setting.
- To explore learning via reward-modulated spike-timing dependent plasticity (STDP) in the IPD.
- To compare STDP with previous reinforcement of stochastic synaptic transmission methods.
Main Methods:
- Developed a computational model with two independent spiking neural networks competing in the IPD game.
- Implemented reward-modulated STDP with eligibility trace for agent learning.
- Extended the learning algorithm with global reinforcement signals to enhance neuronal-level competition.
Main Results:
- The system successfully exhibited mutual cooperation between agents in the IPD game.
- Cooperative outcomes were achieved within a relatively short learning period, enhancing reward accumulation.
- Learning and cooperative outcomes were further enhanced by strong agent memory (high eligibility trace time constant) and firing irregularity (partial somatic reset).
Conclusions:
- Reward-modulated STDP is an effective learning mechanism for achieving cooperation in multiagent systems like the IPD.
- The proposed model demonstrates that SNNs can learn complex cooperative behaviors through biologically inspired plasticity rules.
- Enhancements like global reinforcement signals and specific neuronal mechanisms improve learning efficiency and cooperative outcomes.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Multi-input and Multi-variable systems
In the absence of...
Comparison between RL and RC circuits
Associative Learning
Classical conditioning, also known...
Observational Learning