人体体低频振荡与强化学习期间的预期值和结果相关
Antoine Collomb-Clerc1, Maëlle C M Gueguen1,2, Lorella Minotti1,3
1Univ. Grenoble Alpes, Inserm, U1216, CHU Grenoble Alpes, Grenoble Institut Neurosciences, 38000, Grenoble, France.
Nature communications
|October 17, 2023
概括
人类的丘脑,特别是前部和后部中部区域,在强化学习中发挥着关键作用. 这些地区的低频振荡在决策过程中追踪预期值和预测错误.
科学领域:
- 神经科学是一个神经科学.
- 认知科学 认知科学
- 计算神经科学是一种神经科学.
背景情况:
- 强化学习依赖于前-三角形电路.
- 泰拉马斯是这些电路的关键,但尚未研究的组成部分.
- 对于thalamic参与人类强化学习的直接电生理学证据是有限的.
研究的目的:
- 为了研究人类 thalamus 在基于强化学习的适应性决策中的作用.
- 在强化学习任务中分析thalamic内电生理学记录.
- 为了识别与预期值和预测错误相关的体内的神经信号.
主要方法:
- 在8名人类参与者中,从前额 Thalamus (ATN) 和背中额 Thalamus (DMTN) 获得了电生理学记录.
- 参与者执行了强化学习任务,包括奖励和惩罚.
- 计算建模用于估计预期值和奖励预测错误.
主要成果:
- 在ATN和DMTN的低频振荡 (LFO,4-12 Hz) 与奖励和惩罚学习期间的估计预期值正相关.
- 甲状腺LFO与结果负相关,表明奖励预测错误的信号.
- 在奖励和惩罚条件之间观察到明显的预测信号,这表明在行动抑制中发挥了作用.
结论:
- 人类的丘脑,特别是ATN和DMTN,直接参与基于强化的决策.
- 塔拉米克LFO编码预期值和预测错误,对于适应性学习至关重要.
- 这些发现阐明了人类 thalamus 中的决策和行动抑制的神经机制.
相关概念视频
Law of Effect
1.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.4K
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Observational Learning
188
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
188
Expected Value
4.0K
The expected value is known as the "long-term" average or mean. This means that over the long term of experimenting over and over, you would expect this average. The expected average is represented by the symbol μ. It is calculated as follows:
4.0K
Reinforcement Schedules
160
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
160
Associative Learning
412
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
412


