密文PVT:金字塔视觉转换器与深度多尺度特征精细化网络,用于密文检测.
My-Tham Dinh1, Deok-Jai Choi1, Guee-Sang Lee1
1Department of Artificial Intelligence Convergence, Chonnam National University, 77 Yongbong-ro, Gwangju 500-757, Republic of Korea.
Sensors (Basel, Switzerland)
|July 14, 2023
概括
我们开发了DenseTextPVT,这是一个有效的方法来检测复杂场景图像中的密集文本. 这种方法提高了文本检测的准确性,并减少了重叠的区域,超过了对基准数据集的现有方法.
科学领域:
- 计算机视觉 计算机视觉
- 机器学习 机器学习
- 图像处理 图像处理
背景情况:
- 在场景图像中检测密集文本是具有挑战性的,因为它具有很高的可变性,复杂性和重叠的文本.
- 现有的方法难以准确识别和细分密集的文本实例.
研究的目的:
- 提出一种高效准确的方法来检测场景图像中的密文.
- 增强特征表示,以更好地检测具有不同特征的文本.
- 为了提高精度,在后处理中减少重叠的文本区域.
主要方法:
- 在多个层面上生成高分辨率功能,以准确检测密文本.
- 设计了深度多尺度特征改进网络 (DMFRN),以增强特征表示.
- 利用像素聚合 (PA) 相似性向量算法将文本像素集群到内核中.
主要成果:
- 提出的DenseTextPVT方法在检测密文本方面表现出有效性.
- 在自然图像中实现了更高的精度和减少重叠的文本区域.
- 在TotalText,CTW1500和ICDAR-2015基准数据集上表现优于现有的方法.
结论:
- 在具有挑战性的场景图像中,DenseTextPVT提供了一种有效的解决方案,用于密集文本检测.
- 该方法能够处理不同的文本尺度,形状和字体,包括小文本,是显著的.
- 该方法有效地解决了目前在密文场景中的方法的局限性.
相关概念视频
Visual System
620
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
620
Vision
53.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.6K
Depth Perception and Spatial Vision
735
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
735


