MPSA-Conformer-CTC/Attention:一种高精度,低复杂度的端到端方法,用于西藏语音识别
Changlin Wu1,2, Huihui Sun1,2, Kaifeng Huang1,2
1School of Mechanical and Electrical Engineering, Huainan Normal University, Huainan 232001, China.
Sensors (Basel, Switzerland)
|November 9, 2024
概括
这项研究使用新型的Conformer-CTC/Attention模型与可能的稀疏注意力和MaxEnt优化来增强西藏语音识别. 该MPSA-Conformer-CTC/Attention模型显著降低了文字错误率和计算需求.
科学领域:
- 人工智能的人工智能
- 语音处理 语音处理
- 计算语言学 计算语言学
背景情况:
- 西藏语音识别面临着精度低,计算成本高的挑战.
- 端到端深度学习网络提供潜在的解决方案,但需要有效的架构和解码策略.
研究的目的:
- 开发一个改进的端到端的西藏语音识别模型.
- 通过建立机制的新型集成来提高准确性和降低计算需求.
主要方法:
- 使用Conformer架构进行特征提取.
- 综合连接式时间分类 (CTC) 和注意机制,用于联合解码.
- 引入了概率论的稀疏注意力,以解决趋同问题.
- 实现了CTC的最大优化,以提高训练稳定性.
主要成果:
- 拟议的MPSA-Conformer-CTC/Attention模型在西藏数据集上实现了显著的文字错误率降低,分别为10.68%和9.57%.
- 与基线模型相比,证明减少了记忆消耗和训练时间.
- 展示了改进的概括能力和整体准确性.
结论:
- 该MPSA-Conformer-CTC/Attention模型代表了西藏语音识别的重大进步.
- 结合了Conformer,CTC,Attention,Probabilistic Sparse Attention和MaxEnt优化,有效地解决了准确性和效率方面的挑战.
更多相关视频
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


