学习视觉能力从演示视频中接地.
IEEE transactions on neural networks and learning systems
|September 11, 2023
概括
本研究引入了一种手动辅助的负担能力接地网络 (HAG-Net),以改善人与物体交互的细分. HAG-Net使用视频中的手指线索准确识别交互区域,优于现有方法.
科学领域:
- 计算机视觉 计算机视觉
- 机器人技术 机器人技术 机器人技术
- 人与计算机的交互
背景情况:
- 视觉负担能力接地识别人与物体之间的交互区域,用于机器人掌握等应用.
- 当前的方法与对象的多个交互可能性和同一区域内的多个人类行为作斗争.
研究的目的:
- 提出一个新的手动支付能力接地网络 (HAG-Net),以解决当前视觉支付能力接地方法的局限性.
- 利用演示视频中的手部位置和动作线索来改进交互区域本地化.
主要方法:
- HAG-Net采用双分支架构,单独处理视频和对象数据.
- 视频分支使用手辅助注意力和长短期记忆 (LSTM) 来进行动作特征聚合.
- 对象分支包含一个语义增强模块 (SEM) 和蒸损失,用于知识传输.
主要成果:
- 对两个数据集的评估表明,HAG-Net在视觉负担能力基础上实现了最先进的性能.
- 该方法有效地消除对象的多重交互可能性,并使用手指线索完善本地化.
结论:
- 通过整合来自演示视频的以手为中心的信息,HAG-Net成功地改善了视觉负担能力接地.
- 拟议的网络提供了一个更强大的解决方案,用于在复杂的场景中对人与物体交互区域进行细分.
相关概念视频
Depth Perception and Spatial Vision
709
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
709
Observational Learning
207
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
207
Gestalt Principles of Perception
340
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
340
Visual Agnosia
234
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
234
Purposive Learning
139
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
139
Perceptual Constancy
437
Perceptual constancy is the ability to recognize that objects remain consistent and unchanged even when their appearance varies due to changes in sensory input. There are four main types of perceptual constancy: size constancy, shape constancy, color constancy, and brightness constancy.
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
437


