通过强化学习建模上下文暗示效应的动态,通过强化学习来建模上下文暗示效应的动态
Yasuhiro Hatori1,2,3, Zheng-Xiong Yuan1,4, Chia-Huei Tseng1,5
1Research Institute of Electrical Communication, Tohoku University, Sendai, Japan.
Journal of vision
|November 19, 2024
概括
这项研究模拟了人类如何学习使用环境背景来更快地搜索对象. 该模型显示,场景背景显著提高了视觉搜索效率,模仿人类学习模式.
科学领域:
- 认知心理学 认知心理学
- 计算神经科学是一种神经科学.
- 计算机视觉 计算机视觉
背景情况:
- 人类利用环境环境来提高对象搜索效率.
- 了解背景提示背后的学习机制对于视觉认知至关重要.
- 之前的模型并没有完全捕捉到视觉搜索中的上下文学习细微差别.
研究的目的:
- 提出和验证视觉搜索中的上下文提示效应的计算模型.
- 调查场景背景如何影响目标本地化学习过程.
- 在受控视觉搜索实验中,将模型性能与人类行为进行比较.
主要方法:
- 开发了一个计算模型,提取全球场景特征来学习目标位置关联.
- 实施了一种学习机制,通过重复暴露来加强场景特征和目标位置之间的关系.
- 进行了两次视觉搜索实验,使用统一的 (字母排列) 和自然场景背景.
主要成果:
- 该模型成功地模拟了与均背景相比,在自然场景中检测目标时观察到的斜动减小.
- 模型的性能复制了人类上下文暗示的关键特征,包括本地学习和刺激新奇性的影响.
- 与统一的背景相比,自然场景表现出优越的上下文暗示效应,与实验发现保持一致.
结论:
- 拟议的模型为理解上下文暗示的计算基础提供了一个可行的框架.
- 场景背景在有效的视觉搜索中发挥着重要作用,因为它促进了学习的关联.
- 这项研究有助于更深入地了解复杂环境中的视觉注意力和学习.
相关概念视频
Reinforcement Schedules
132
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
132
Timing and Consequences on Behavior
82
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
82
Cognitive Learning
220
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
220
Generalization, Discrimination, and Extinction
450
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
450
Reinforcement
180
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
180
Real-World Application of Classical Conditioning
527
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
527


