联合变化和ZhuYin数据集用于传统中文文档增强
Shi-Wei Lo1, Hsiu-Mei Chou2, Jyh-Horng Wu2
1National Center for High-Performance Computing, Hsinchu, Taiwan. LSW@narlabs.org.tw.
Scientific data
|November 28, 2024
概括
一个新的数据集,联合变化和朱 (JVZY),解决了传统中文文档增强培训数据的稀缺问题. 它具有20,000张具有多种降解的图像,有助于人工智能开发.
科学领域:
- 计算机科学 计算机科学
- 人工智能的人工智能
- 数字图像处理 数字图像处理
背景情况:
- 数字文档质量对于信息管理至关重要,但往往因注释,扭曲和污染而退化.
- 深度学习方法增强了文档,但需要广泛的,高质量的数据集进行培训和评估.
- 现有的基准数据集很少,特别是对于传统中国文献,这阻碍了这一领域的进步.
研究的目的:
- 引入一个新的,大规模的数据集,用于传统中文文档增强.
- 提供一个资源,解决传统中国语音符号和文档退化所面临的具体挑战.
- 促进先进的人工智能应用程序的开发,以增强退化的传统中文文档.
主要方法:
- 创建联合变异和朱 (JVZY) 数据集.
- 包括2万张图像和192万个单词.
- 整合了各种文档退化特征和独特的传统中文语音符号.
主要成果:
- JVZY数据集为传统中文文档增强提供了全面的资源.
- 它包含了广泛的降解类型和特定的语言特征.
- 数据集被设计成一个不断变化的资源.
结论:
- JVZY数据集填补了传统中文文档增强资源的关键缺口.
- 它将加速对该领域人工智能模型的研究和开发.
- 此资源支持创建更有效的文档增强应用程序.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Variance
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.The standard deviation measures the spread in the same units as the data.


