新奇的多巴胺编码促进了高效的不确定性驱动的探索.
Yuhao Wang1, Armin Lak2, Sanjay G Manohar3
1MRC Brain Network Dynamics Unit, University of Oxford, Oxford, United Kingdom.
PLoS computational biology
|April 16, 2024
概括
动物通过评估奖励的不确定性来探索新的环境. 一个新的基底腺模型,使用多巴胺作为新信号,推动了高效的,基于不确定性的探索,优于其他策略.
科学领域:
- 神经科学是一个神经科学.
- 计算神经科学是一种神经科学.
- 强化学习是一种强化学习.
背景情况:
- 动物必须探索不熟悉的环境,以学习动作奖励关联,并优化决策.
- 有效的探索需要评估和利用行动奖励知识中的不确定性.
研究的目的:
- 为基于不确定性驱动的探索提出基底的新型计算模型.
- 调查条形状通路和多巴胺类神经元在估计奖励不确定性和新奇性信号中的作用.
主要方法:
- 开发了基底功能的计算模型,整合了直接和间接的条状通路,用于奖励平均值和差异估计.
- 利用电生理学数据验证了基底状腺模型.
- 从神经模型获得的探索策略与行为数据相匹配.
- 在模拟中比较模型启发的探索策略与经典算法,如UCB.
主要成果:
- 拟议的基底腺模型有效地估计了奖励分布的平均值和差异.
- 电生理学数据支持该模型对基底腺功能的表示.
- 与UCB相比,灵感来自基底性腺模型的策略在模拟中表现出优越的性能.
- 模型与行为数据相匹配的结果与理想化的规范模型相美.
结论:
- 在基底中编码新奇性的短暂多巴胺信号有助于不确定性表示.
- 这种不确定性表示有效地推动了强化学习的探索.
- 该模型提供了一个生物学上可信的机制,用于在不确定性下进行最佳决策.
相关概念视频
Uncertainty: Overview
552
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
552
The Availability Heuristic
6.0K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
6.0K
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
The Anchoring-and-Adjustment Heuristic
7.2K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.2K
Generalization, Discrimination, and Extinction
545
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
545


