Eligibility-trace-gated depression in predecessor feature learning enables post-reward exploration
Incheol Seo1,2,3, Sun-Hyun Park4, Hyunsu Lee5,6,7
1Department of Immunology, School of Medicine, Kyungpook National University, Gukchaebosang-Ro, Daegu, 41944 Republic of Korea.
Abstract:
Spatial navigation and exploration require flexible learning mechanisms that can adapt to changing environmental demands. While dopamine's role in reward prediction error is well-established, the computational functions of other neuromodulators in reinforcement learning remain less understood. Here, we investigate how acetylcholine (ACh) modulation affects predecessor feature (PF)-based learning, a computational framework that combines successor representations with eligibility traces for retrospective credit assignment. We developed an ACh-modulated PF algorithm (ACh-PF) implementing synaptic depression via eligibility-trace outer products ([Formula: see text]), hypothesized to promote exploration by attenuating reinforcement of recently traversed transitions. Using n-arm radial mazes, we compared conventional navigation (episodes terminate at reward) with a post-reward exploration criterion requiring visits to all arm endpoints. In conventional mode, all agents achieved near-optimal performance. Under the post-reward criterion, the PF baseline largely failed, whereas ACh-PF exhibited a non-monotonic dependence on [Formula: see text]: performance improved sharply within a narrow intermediate regime and degraded at higher gains, consistent with a knee-like transition followed by over-depression. The effective window narrowed with increasing spatial and action-space complexity; short two-arm mazes ([Formula: see text] - 6) supported near-ceiling reward across a broader range, whereas longer arms ([Formula: see text]) required tighter tuning and showed reduced efficiency. In multi-arm mazes, efficient exploration persisted only in the least demanding conditions (e.g., [Formula: see text]), collapsing toward timeouts as arm length and arm number increased. These results link cholinergic-like synaptic depression to flexible exploration while revealing scaling limits in complex environments.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s11571-026-10477-5.
More Related Videos
06:46Automated, Long-term Behavioral Assay for Cognitive Functions in Multiple Genetic Models of Alzheimer's Disease, Using IntelliCage
Published on: August 4, 2018
08:05A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Related Concept Videos
Long-term Depression
Calcium Ion Concentration Mechanism
If over time, all...
Long-term Depression
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Hindsight Biases
