Related Experiment Video
Updated: Oct 6, 2025

Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
A nonlinear hidden layer enables actor-critic agents to learn multiple paired association navigation.
M Ganesh Kumar1,2,3, Cheston Tan4, Camilo Libedinsky1,2,5,6
1Integrative Sciences and Engineering Programme, NUS Graduate School, National University of Singapore, Singapore 119077, Singapore.
Biologically plausible actor-critic agents struggle with multiple reward locations. A novel agent using a feedforward or recurrent network with temporal difference learning successfully navigates complex cued navigation tasks.
Area of Science:
- Neuroscience and computational modeling
- Rodent learning and navigation
Background:
- Multiple cued reward location tasks are vital for studying rodent learning.
- Deep reinforcement learning agents can learn these tasks but lack biological plausibility.
- Classic biologically plausible actor-critic agents are limited to single reward locations.
Purpose of the Study:
- To investigate biologically plausible agents for multiple cue-reward location navigation.
- To identify limitations of classic actor-critic agents in complex navigation tasks.
- To develop and test novel computational models for enhanced rodent learning.
Main Methods:
- Computational modeling of actor-critic agents.
- Adaptation of agents to single reward locations and displacements.
- Introduction of a feedforward nonlinear hidden layer processing place cell and cue information.
- Implementation of temporal difference error-modulated plasticity.
- Comparison with a recurrent reservoir network architecture.
Main Results:
- Classic actor-critic agents navigated single reward locations but failed with multiple associations.
- A novel agent with a feedforward hidden layer successfully learned multiple cue-reward navigation.
- Replacing the feedforward layer with a recurrent reservoir network led to faster learning.
Conclusions:
- Classic biologically plausible agents are insufficient for complex multi-cue navigation tasks.
- A novel computational approach incorporating a hidden layer with TD error-modulated plasticity enhances agent capabilities.
- Recurrent networks offer further improvements in learning speed for these navigation tasks.
Related Concept Videos
Associative Learning
Classical conditioning, also known...
Observational Learning
Purposive Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Actor-Observer Effect
Real-World Application of Classical Conditioning
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...

