可以解释的基于连接主义的时间分类的场景文本识别
Rina Buoy1, Masakazu Iwamura1, Sovila Srun2
1Department of Core Informatics, Graduate School of Informatics, Osaka Metropolitan University, Osaka 599-8531, Japan.
Journal of imaging
|November 24, 2023
概括
本研究引入了一种用于场景文本识别 (STR) 的新方法,该方法将连接主义时间分类 (CTC) 的效率与改进的可解释性相结合. 该方法提高了STR模型中的字符定位和预测透明度.
科学领域:
- 计算机视觉 计算机视觉
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 连接主义时间分类 (CTC) 由于其效率在场景文本识别 (STR) 中被广泛使用,但由于其依赖于1D序列,缺乏可解释性.
- 现有的基于2D注意力的方法提供了更好的准确性和本地化,但在计算上是密集的,带来了延迟挑战.
研究的目的:
- 开发一种低延迟的STR方法,通过字符定位提供模型可解释性.
- 通过使它们能够处理2D空间信息以提高可解释性来增强1D CTC解码器.
主要方法:
- 提出一种基于边缘化的方法来处理二维特征图,预测高度和类维度的联合概率分布.
- 引入了一个"关联地图"用于字符定位和解释,在基于注意力的模型中扮演着类似于交叉注意力地图的角色.
- 视觉变压器 (ViT) 与1D CTC解码器 (ViT-CTC) 架构相结合,用于STR.
主要成果:
- 在对基准的识别准确性方面,ViT-CTC模型的表现优于基于CTC的最先进 (SOTA) 方法.
- 与基线变压器解码器模型相比,ViT-CTC模型的速度提升高达12倍,精度降低最小.
- 从关联图估计的字符位置显示与地面真实界限框和交叉注意力图有很强的对齐.
结论:
- 提出的基于边缘化的方法成功地将字符本地化和可解释性集成到STR的1D CTC解码器中.
- 对于场景文本识别任务,ViT-CTC提供了高识别准确度,低延迟和增强的模型解释性的令人信服的平衡.
相关概念视频
Force Classification
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Systems-I
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:


