Related Experiment Video
Updated: May 24, 2025

08:08
Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
Published on: April 1, 2016
9.3K
CCDPlus: Towards Accurate Character to Character Distillation for Text Recognition
Summary
CCDPlus enhances scene text recognition by using unlabeled real data (URD) and labeled synthetic data (LSD) through a character-to-character distillation method. This approach overcomes domain gaps and improves accuracy for real-world text recognition tasks.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Scene text recognition (STR) methods often rely on large-scale synthetic data (LSD), but a synth-to-real domain gap limits performance.
- Unlabeled real data (URD) holds valuable information, yet effectively utilizing it for STR remains challenging.
- Previous self-supervised methods on URD face issues like coarse representations, inflexible augmentation, and real-to-synth domain drift.
Purpose of the Study:
- To propose CCDPlus, a novel character-to-character distillation method for robust scene text recognition.
- To address the limitations of existing methods by integrating URD and LSD within a unified framework.
- To improve the efficiency and robustness of STR models in real-world scenarios.
Main Methods:
- CCDPlus employs a joint supervised and self-supervised learning framework for scene text recognition.
- It delineates fine-grained character structures on URD as representation units by transferring knowledge from LSD online.
- The method enables flexible character-to-character distillation with versatile data augmentation, extracting generalizable character-level features.
Main Results:
- CCDPlus achieves state-of-the-art (SOTA) performance, outperforming supervised, semi-supervised, and self-supervised methods.
- It demonstrates an average improvement of 1.8% over supervised, 0.6% over semi-supervised, and 1.1% over self-supervised methods on standard datasets.
- Significant improvement of 6.1% is observed on the challenging Union14M-L dataset.
Conclusions:
- CCDPlus effectively bridges the synth-to-real domain gap in scene text recognition.
- The character-to-character distillation approach enhances feature representation and recognition accuracy.
- The unified framework successfully combines self-supervised learning on URD with supervised learning on LSD for superior performance.

