SEG-SLAM:动态室内RGB-D视觉SLAM集成几何和YOLOv5基础的语义信息
Peichao Cong1, Jiaxing Li1, Junjie Liu1
1School of Mechanical and Automotive Engineering, Guangxi University of Science and Technology, Liuzhou 545006, China.
Sensors (Basel, Switzerland)
|April 13, 2024
概括
本研究介绍了SEG-SLAM,这是一个动态视觉SLAM系统,可以提高机器人导航的准确性. 它有效地处理动态对象,在现实世界的场景中改进本地化和映射.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 计算机视觉 计算机视觉
- 人工智能的人工智能
背景情况:
- 同时定位和映射 (SLAM) 对移动机器人至关重要.
- 现有的视觉SLAM系统由于移动的物体而与动态环境作斗争.
- 动态对象显著降低了传统SLAM的准确性和稳定性.
研究的目的:
- 为改进机器人导航提出一种新的动态视觉SLAM系统 (SEG-SLAM).
- 通过解决动态对象带来的挑战来提高视觉SLAM的性能.
- 为了在复杂的环境中实现更准确,更强大的本地化和映射.
主要方法:
- SEG-SLAM将ORB-SLAM3框架与YOLOv5深度学习模型集成在一起,用于对象检测和语义细分.
- 融合模块使用先前和深度信息识别动态对象.
- 差异化的特征点拒绝策略和极几何学被用于处理动态对象.
主要成果:
- 该系统有效地识别和提取有关动态对象的信息.
- 基于对象类型和环境数据,动态特征点被战略性地拒绝.
- 通过将拒绝结果与深度信息融合,生成一个静态密集的3D地图.
结论:
- 与现有的动态视觉SLAM算法相比,SEG-SLAM显著提高了本地化和映射精度.
- 拟议的方法在现实场景和公共数据集中表现出增强的稳定性.
- 在具有动态元素的环境中,SEG-SLAM为移动机器人导航提供了更可靠的解决方案.
更多相关视频
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.0K
07:09Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
Published on: May 2, 2019
6.1K
相关概念视频
Depth Perception and Spatial Vision
643
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
643
Color Vision
567
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
567
Light Acquisition
8.5K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.5K
Visual System
579
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
579
Convolution: Math, Graphics, and Discrete Signals
249
In any LTI (Linear Time-Invariant) system, the convolution of two signals is denoted using a convolution operator, assuming all initial conditions are zero. The convolution integral can be divided into two parts: the zero-input or natural response and the zero-state or forced response, with t0 indicating the initial time.
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
249
Vision
53.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.2K
