SPTS v2:单点场景文本发现
IEEE transactions on pattern analysis and machine intelligence
|September 5, 2023
概括
我们的新框架,SPTS v2,只使用单点注释实现了高性能场景文本发现. 这种方法显著降低了注释成本,并且比以前的方法推断速度快19倍.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 端到端的场景文本发现集成了文本检测和识别.
- 传统的方法需要昂贵的手动注释,如边界框或多边形.
- 单点注释提供了一个更具成本效益的替代方案.
研究的目的:
- 开发一个新的框架,SPTS v2,用于使用单点注释有效地发现场景文本.
- 通过专用解码器利用变压器架构来提高性能和速度.
- 展示场景文本识别中单点表示的可行性和优势.
主要方法:
- SPTS v2 使用自动回归变压器与实例分配解码器 (IAD) 进行顺序中心点预测.
- 一个并行识别解码器 (PRD) 同时处理文本识别,减少序列长度要求.
- 两个解码器共享参数,并通过有效的信息传输过程进行交互.
主要成果:
- SPTS v2 在使用单点注释的基准数据集上实现了最先进的性能.
- 该框架在较少的参数和19倍更快的推断速度方面表现出卓越的效率,与以前的方法相比.
- 实验结果表明,在场景文本发现中,人们可能更喜欢单点表示.
结论:
- SPTS v2 在场景文本识别中提供了显著的进步,通过在最小的注释努力下实现高性能.
- 拟议的方法为现实世界的场景文本发现应用提供了更实用和可扩展的解决方案.
- 这项工作为现场文本发现开辟了超越当前范式的新途径,强调了单点注释的有效性.
更多相关视频
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
15.8K
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
9.1K
相关概念视频
Fixing Double-strand Breaks
3.1K
3.1K
Focusing of Light in the Eye
2.9K
Light rays enter the eye through the cornea, a transparent dome-shaped tissue that is the eye's outermost layer. The cornea bends or refracts, light rays traveling to the pupil. The shape of the cornea determines how much of the light is bent and whether the image will be focused correctly on the retina at the back of the eye. Once the light has passed through both refraction layers, it converges into a single focal point onto a small area. This is where photoreceptors start transforming...
2.9K
Detection of Gross Error: The Q Test
6.2K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.2K
Proofreading
54.2K
Overview
54.2K
Tip-of-the-Tongue Phenomenon
167
The tip-of-the-tongue (TOT) phenomenon is a cognitive experience characterized by a temporary inability to retrieve specific information from memory despite having a strong feeling of knowing the information. Although individuals cannot access the target word or detail, they frequently recall related elements, such as its initial letter, syllable count, or context. This partial retrieval often causes frustration, as one might recognize a familiar face or know that a name starts with a specific...
167
Additional Subnuclear Structures
2.1K
2.1K
