相关实验视频
Updated: Jul 1, 2025

09:45
New Variations for Strategy Set-shifting in the Rat
Published on: January 23, 2017
8.2K
条形动作选择和强化学习的动态
Jack Lindsey1, Jeffrey E Markowitz2, Winthrop F Gillis3
1Kavli Institute for Brain Science, Columbia University, New York, NY, USA.
bioRxiv : the preprint server for biology
|March 11, 2024
概括
强化学习模型面临着棘手投射神经元 (SPN) 可塑性的挑战. 同时激活直接和间接途径的SPNs解决了这一问题,使基底腺节能够有效地学习.
科学领域:
- 神经科学是一个神经科学.
- 计算神经科学是一种神经科学.
- 系统神经科学 系统神经科学
背景情况:
- 脊椎条纹体中的棘状投射神经元 (SPN) 是基底质中强化学习模型的核心.
- 现有的模型与已知的SPN突触可塑性规则存在不一致.
研究的目的:
- 解决条形强化学习模型和SPN突触可塑性之间的不一致性.
- 提出一个修订后的模型学习和行动选择在基底.
主要方法:
- 对直接通路 (dSPN) 和间接通路 (iSPN) 神经元的突触可塑性规则的分析.
- 强化学习的计算建模,包括SPN函数.
- 使用条纹记录进行验证.
主要成果:
- 目前模拟的iSPN可塑性通过强化负面结果来阻碍学习.
- 通过不同的输入同时激活功能上对立的dSPN和iSPN,可以逆转这种病理效应.
- 这解决了条形学习模型和SPN可塑性之间的冲突.
结论:
- 一个修订后的模型允许在没有干扰的情况下进行学习和行动选择信号的多重化.
- 这个框架支持超越标准时间差模型的学习算法.
- 这些发现提供了SPN函数在强化学习中的更准确的表示.
相关概念视频
Law of Effect
1.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.4K
Instinctive Drift
218
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
218
Timing and Consequences on Behavior
90
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
90
Operant Conditioning
1.6K
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
1.6K
Reinforcement Schedules
145
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
145
Generalization, Discrimination, and Extinction
552
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
552

