适应多语言视觉语言转换器用于低资源的乌尔都语光学字符识别 (OCR)
Musa Dildar Ahmed Cheema1, Mohammad Daniyal Shaiq1, Farhaan Mirza2
1Department of Artificial Intelligence and Data Science, National University of Computer and Emerging Sciences, Islamabad, Pakistan.
PeerJ. Computer science
|May 3, 2024
概括
本研究介绍了ViLanOCR,这是一个双语的乌尔都语和英语的光学字符识别 (OCR) 系统. 它使用先进的变压器模型来实现低资源语言数字化的最新结果.
科学领域:
- 自然语言处理自然语言处理.
- 计算机视觉 计算机视觉
- 数字人文学科 数字人文学科
背景情况:
- 低资源语言对准确的光学字符识别 (OCR) 提出了重大挑战.
- 现有的OCR系统经常与资源不足的语言的语言复杂性作斗争.
- 这些语言的书面内容数字化需要专门的方法.
研究的目的:
- 推出ViLanOCR,一种乌尔都语和英语的创新双语OCR系统.
- 解决OCR在资源较少的语言环境中的特定挑战.
- 与现有的OCR解决方案相比,为了证明卓越的性能.
主要方法:
- 开发ViLanOCR,一个双语OCR系统.
- 利用基于多语言变压器的先进语言模型.
- 在乌尔都语UHWR数据集上使用字符错误率 (CER) 度量进行评估.
主要成果:
- 在乌尔都语UHWR数据集上,ViLanOCR实现了1.1%的字符错误率 (CER).
- 该系统展示了乌尔都文手写数字化的最新性能.
- 实验结果证实了拟议方法的有效性.
结论:
- ViLanOCR为乌尔都语等低资源语言的OCR提供了一个强大的解决方案.
- 先进的变压器模型在具有挑战性的语言环境中有效提高OCR准确性.
- 该系统超越了乌尔都文手写数字化的当前最先进的基线.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


