从端到端的语义场景中进行视觉位置识别 文本特征 场景特征
1Electrical Engineering Department, Chabahar Maritime University, Chabahar, Iran.
Frontiers in robotics and AI
|October 1, 2024
概括
机器人现在可以更好地使用图像中的文本来识别地点. 我们的新方法有效地检测和读取文本,即使它是.
科学领域:
- 计算机视觉 计算机视觉
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
背景情况:
- 城市环境提供了丰富的视觉文本线索.
- 机器人可以利用这些文本功能来增强位置识别.
- 目前的视觉位置识别 (VPR) 方法在文本变化方面面临挑战.
研究的目的:
- 引入一种新的技术,让机器人利用城市文本进行视觉位置识别 (VPR).
- 通过整合场景文本检测和识别来提高机器人本地化和映射能力.
主要方法:
- 开发了一种针对VPR的端到端场景文本检测和识别技术.
- 该模型捕获文本字符串和边界框,解决诸如任意形状,照明变化和遮蔽等挑战.
- 利用端到端的场景文本发现框架,在各种条件下进行强大的文本捕获.
主要成果:
- 在自收集的TextPlace (SCTP) 基准数据集上进行了实验性评估.
- 与最先进的方法相比,提出的方法显示出更高的性能.
- 在VPR任务的精度和回忆方面取得了显著的改进.
结论:
- 拟议的端到端场景文本发现框架对VPR有效.
- 这种方法显示了改善复杂的城市环境中的机器人本地化和绘制的巨大潜力.
- 该方法成功地应对了不规则和封闭文本所带来的挑战.
更多相关视频
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
8.9K
05:38Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
2.3K
相关概念视频
Vision
53.0K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.0K
Visual System
553
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
553
Depth Perception and Spatial Vision
602
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
602
Parallel Processing
145
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
145
