使用随机释放可塑性的奖励优化学习
Yuhao Sun1,2, Wantong Liao1,2, Jinhao Li1,3
1Laboratory of Brain and Intelligence, Tsinghua University, Beijing, China.
Frontiers in neural circuits
|September 2, 2025
概括
我们介绍了奖励优化的随机释放可塑性 (RSRP), RSRP实现了强大而有效的奖励驱动的学习,与人工智能和神经科学中的既定方法相美.
科学领域:
- 计算神经科学
- 人工智能
- 机器学习
背景情况:
- 突触可塑性使神经系统的适应性学习成为一种奖励驱动学习的生物学可信模型.
- 一个关键的挑战是制定与错误反向传播相匹配的可塑性规则.
研究的目的:
- 引入奖励优化的随机释放可塑性 (RSRP),这是一个新的学习框架.
- 使用自然梯度估计来最大化奖励信号的可塑性规则.
- 评估RSRP在强化学习和数字分类任务中的性能和稳定性.
主要方法:
- 在RSRP框架内模型突触释放为参数分布.
- 使用自然梯度估计来推导RSRP学习规则.
- 在生物可信的神经网络中验证RSRP并与近接政策优化 (PPO) 和错误反向传播进行比较.
主要成果:
- 在强化学习方面,RSRP表现出具有竞争力的表现和稳定性,与PPO相当.
- 在数字分类任务中,RSRP的准确性与错误反向传播相当.
- 奖励规范化被认为是稳定RSRP的关键机制.
结论:
- RSRP提供了一个强大而有效的突触可塑性学习规则.
- 这些发现对人工智能和实验神经科学都有影响,特别是在不连续的强化学习场景中.
相关概念视频
Neuroplasticity
752
Neuroplasticity reflects the brain's remarkable capacity to adapt and evolve, responding dynamically to learning, experiences, or injury by reorganizing its neural circuitry. This reorganization involves creating new neural connections and refining old ones through a series of biological processes that contribute to the brain's lifelong development and adaptability.
752
Long-term Potentiation
2.9K
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Hebbian LTP
LTP can occur when...
Hebbian LTP
LTP can occur when...
2.9K
Plasticity
2.5K
Plasticity is the property where an object loses its elasticity and undergoes irreversible deformation, even after the deformation forces are eliminated. If a material deforms irreversibly without increasing stress or load, then this is called ideal plasticity. For example, when a force is applied to an aluminum rod, it changes its shape, but it does not return to its original shape once the force is removed. Plastic deformation or ductility is thus a permanent deformation or change in the...
2.5K
Purposive Learning
204
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
204
Long-term Depression
2.6K
Long-term depression, or LTD, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTD is the process of synaptic weakening that occurs over time between pre and postsynaptic neuronal connections. The synaptic weakening of LTD works in opposition to synaptic strengthening by long-term potentiation (LTP) and together are the main mechanisms that underlie learning and memory.
Calcium Ion Concentration Mechanism
If over...
Calcium Ion Concentration Mechanism
If over...
2.6K


