EgoVision是一个YOLO-ViT混合体,用于强大的自我中心物体识别
Umm E Sadima1, Yazeed Alkharijah2, Danish Hamid1
1Department of Creative Technologies, Faculty of Computing and Artificial Intelligence (FCAI), Air University, Islamabad, 44000, Pakistan.
Scientific reports
|October 6, 2025
概括
混合深度学习框架EgoVision通过结合YOLOv8和视觉转换器来增强自我中心的对象识别. 这种轻量级模型在机器人和增强现实领域的实时应用中实现了高精度.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
背景情况:
- 自我中心的视觉,或第一人称的视角,对于辅助技术,增强现实和人机交互至关重要.
- 在以自我为中心的视觉中,对象识别面临着诸如阻塞,运动模糊和视角变化等挑战.
- 现有的方法与实时自我中心物体识别的计算需求作斗争.
研究的目的:
- 介绍EgoVision,一种新的混合深度学习框架,用于在静态的自我中心框架中对象分类.
- 将YOLOv8的空间精度与视觉转换器 (ViT) 的全球上下文推理融合在一起.
- 使用HOI4D数据集,为机器人和增强现实中的应用实现实时对象识别.
主要方法:
- 开发了EgoVision,这是一个轻量级的混合深度学习框架,结合了YOLOv8和视觉转换器 (ViT).
- 采用关键提取策略和特征金字塔网络,以高效地处理多尺度的时空特征.
- 利用HOI4D数据集在自我中心的框架中训练和评估静态对象识别.
主要成果:
- 对于复杂的对象类,如"水"和"椅子",EgoVision的准确度高达99%.
- 与现有模型相比,该框架在多个指标上表现出卓越的表现.
- EgoVision保持了高效率,适合在可穿戴设备和边缘设备上部署.
结论:
- EgoVision代表了自我中心物体识别的重大进步,提供了高精度和效率.
- 混合架构有效地解决了第一人称视角对象分类的挑战.
- 在实时应用中,EgoVision为下一代自我中心的人工智能系统提供了坚实的基础.
相关概念视频
Vision
59.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.3K
Light Acquisition
9.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.4K
Force Classification
2.3K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
2.3K
Observational Learning
832
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
832
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Deconvolution
543
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
543


