Related Experiment Video
Updated: May 11, 2026

09:27
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
9.9K
Joint variation and ZhuYin dataset for Traditional Chinese document enhancement
Shi-Wei Lo1, Hsiu-Mei Chou2, Jyh-Horng Wu2
1National Center for High-Performance Computing, Hsinchu, Taiwan. LSW@narlabs.org.tw.
Scientific Data
|November 28, 2024
Summary
A new dataset, Joint Variation and ZhuYin (JVZY), addresses the scarcity of training data for Traditional Chinese document enhancement. It features 20,000 images with diverse degradation, aiding AI development.
Area of Science:
- Computer Science
- Artificial Intelligence
- Digital Image Processing
Background:
- Digital document quality is crucial for information management but often degraded by annotations, distortion, and stains.
- Deep learning methods enhance documents but require extensive, high-quality datasets for training and evaluation.
- Existing benchmark datasets are scarce, especially for Traditional Chinese documents, hindering advancement in this area.
Purpose of the Study:
- To introduce a novel, large-scale dataset for Traditional Chinese document enhancement.
- To provide a resource that addresses the specific challenges of Traditional Chinese phonetic symbols and document degradation.
- To facilitate the development of advanced AI applications for enhancing degraded Traditional Chinese documents.
Main Methods:
- Creation of the Joint Variation and ZhuYin (JVZY) dataset.
- Inclusion of 20,000 images and 1.92 million words.
- Incorporation of diverse document degradation characteristics and unique Traditional Chinese phonetic symbols.
Main Results:
- The JVZY dataset offers a comprehensive resource for Traditional Chinese document enhancement.
- It contains a wide range of degradation types and specific linguistic features.
- The dataset is designed as a continuously evolving resource.
Conclusions:
- The JVZY dataset fills a critical gap in resources for Traditional Chinese document enhancement.
- It will accelerate research and development of AI models for this domain.
- This resource supports the creation of more effective document enhancement applications.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Variance
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.The standard deviation measures the spread in the same units as the data.

