通过双流交互神经架构的连续视觉语音识别
Hao Yuan1, Yakun Zhang2, Xingyu Zhang2
1School of Advanced Manufacturing and Robotics, Peking University, 100871, Beijing, China; Defense Innovation Institute, Academy of Military Sciences, 100071, Beijing, China; Intelligent Game and Decision Laboratory, 100071, Beijing, China.
这项研究引入了一种新的双流结构,用于唇阅读,该架构通过使用顺序视觉知识来增强细粒度的视觉特征提取. 该方法在多个数据集上取得了最先进的结果,提高了视觉语音识别的准确性和稳定性.
科学领域:
- 计算机科学 计算机科学
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 当前深度神经网络的唇读方法在细粒度特征提取方面扎,影响了关节细节的捕获.
- 保护整体语义通常会损害唇读模型中关键的局部视觉特征的提取.
研究的目的:
- 引入一种新的范式,用于句子级口唇阅读,使用顺序的视觉知识和双流架构.
- 解决现有的唇读模型中细粒度特征提取的局限性.
- 为了提高视觉语音识别系统的准确性和稳定性.
主要方法:
- 连续视觉知识的概念化和双流架构的开发.
- 整合了顺序的视体动态,以增强和细分的注意力.
- 通过多个途径促进角色和视觉细分信息之间的相互作用.
主要成果:
- 在多个句子级别的唇读数据集 (GRID,CMLR,LRS2,LRS3) 上实现了最先进的性能.
- 在GRID上,文字错误率 (WER) 低至0.6%,在CMLR上,字符错误率 (CER) 低至9.9%.
- 展示了对大规模预培训和数据腐败的稳定性,突出了视觉知识的普遍性.
结论:
- 拟议的双流架构有效地解决了唇阅读中的细粒度特征保存问题.
- 顺序的视觉语言知识对于弥合跨语言差距和增强视觉语音识别至关重要.
- 该方法比目前的高性能模型具有显著优势,具有进一步研究和应用的潜力.
更多相关视频
10:25Dual-color Correlative Light and Electron Microscopy for the Visualization of Interactions between Mitochondria and Lysosomes
Published on: September 27, 2024
08:32Ultrasound Images of the Tongue: A Tutorial for Assessment and Remediation of Speech Sound Errors
Published on: January 3, 2017
相关概念视频
Stream Function
Steady Flow of a Fluid Stream
During this process, the momentum of the fluid within the control volume remains constant over the time interval dt. By applying the...
Polymer Classification: Architecture
Neural Regulation
ATP Driven Pumps I: An Overview
There are four main types of ATP-driven pumps - P-type, V-type, F-type, and ABC transporter. All these pumps are of varying complexities and...
Predator-Prey Interactions
