Eligibility-trace-gated depression in predecessor feature learning enables post-reward exploration
Incheol Seo1,2,3, Sun-Hyun Park4, Hyunsu Lee5,6,7
1Department of Immunology, School of Medicine, Kyungpook National University, Gukchaebosang-Ro, Daegu, 41944 Republic of Korea.
Acetylcholine (ACh) modulation enhances spatial exploration in reinforcement learning by employing synaptic depression. This ACh-modulated predecessor feature (PF) algorithm shows improved performance in complex mazes, but with scaling limitations.
Area of Science:
- Neuroscience
- Computational Neuroscience
- Reinforcement Learning
Background:
- Dopamine's role in reward prediction error is established in reinforcement learning.
- Computational functions of other neuromodulators, like acetylcholine (ACh), are less understood.
- Flexible learning and exploration are crucial for adapting to changing environments.
Purpose of the Study:
- Investigate the effect of ACh modulation on predecessor feature (PF)-based learning.
- Develop an ACh-modulated PF algorithm (ACh-PF) using synaptic depression to promote exploration.
- Compare ACh-PF performance against a baseline PF algorithm in varying maze complexities.
Main Methods:
- Developed an ACh-modulated PF algorithm (ACh-PF) implementing synaptic depression.
- Utilized eligibility-trace outer products for synaptic depression in the ACh-PF model.
- Tested algorithms in n-arm radial mazes with conventional and post-reward exploration criteria.
Main Results:
- ACh-PF demonstrated improved performance under a post-reward exploration criterion compared to baseline PF.
- ACh-PF performance showed a non-monotonic dependence on the synaptic depression parameter, with an optimal intermediate regime.
- Performance degraded with increased spatial and action-space complexity, revealing scaling limits.
Conclusions:
- Cholinergic-like synaptic depression is linked to flexible exploration in spatial navigation.
- The ACh-PF algorithm offers a computational framework for understanding ACh's role in exploration.
- Environmental complexity imposes scaling limits on the efficiency of exploration strategies.
More Related Videos
06:46Automated, Long-term Behavioral Assay for Cognitive Functions in Multiple Genetic Models of Alzheimer's Disease, Using IntelliCage
Published on: August 4, 2018
08:05A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Related Concept Videos
Long-term Depression
Calcium Ion Concentration Mechanism
If over time, all...
Long-term Depression
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Hindsight Biases
