Related Experiment Video
Updated: Jun 25, 2026

Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
A spiking neural network model of an actor-critic learning agent
Wiebke Potjans1, Abigail Morrison, Markus Diesmann
1Computational Neuroscience Group, RIKEN Brain Science Institute, Wako City, Saitama 351-0198, Japan. wiebke_potjans@brain.riken.jp
This study introduces a spiking neural network model for reinforcement learning, demonstrating its ability to adapt behavior for maximizing rewards. The model effectively learns complex tasks, bridging the gap between computational algorithms and neural processes.
Area of Science:
- Computational neuroscience
- Artificial intelligence
- Behavioral adaptation
Background:
- Behavioral adaptation is essential for survival, with reinforcement learning offering effective strategies.
- The neural compatibility of temporal-difference learning algorithms remains an open question.
Purpose of the Study:
- To present a spiking neural network model implementing actor-critic temporal-difference learning.
- To investigate the neural plausibility of reinforcement learning algorithms.
Main Methods:
- Developed a spiking neural network model combining local plasticity rules and a global reward signal.
- Utilized a gridworld task with sparse rewards to test the network's learning capabilities.
- Derived a quantitative mapping between network parameters and algorithmic variables.
Main Results:
- The model successfully solved a nontrivial gridworld task with sparse rewards.
- Demonstrated that the spiking neural network learns at a comparable speed to discrete-time algorithms.
- Achieved equivalent equilibrium performance compared to standard reinforcement learning formulations.
Conclusions:
- The proposed spiking neural network model provides a biologically plausible implementation of actor-critic temporal-difference learning.
- The findings suggest a strong compatibility between temporal-difference learning and neural computation.
- The model offers a framework for understanding adaptive behavior in biological and artificial systems.
Related Concept Videos
Actor-Observer Effect
Observational Learning
Neural Regulation
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Multi-input and Multi-variable systems
In the absence of...
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
