Related Experiment Video
Updated: Sep 16, 2025

08:08
Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
Published on: April 1, 2016
9.4K
A syllable-character collaborative model for enhanced Pinyin and Chinese recognition
Zeyuan Chen1, Cheng Zhong1,2, Danyang Chen1,2
1School of Computer, Electronics and Information, Guangxi University, Nanning, Guangxi, China.
Plos One
|July 7, 2025
Summary
Chinese speech recognition models struggle with complex pronunciation. Our Syllable-Character Collaborative Model (SCCM) improves accuracy by incorporating phonetic elements like pinyin, significantly reducing character errors.
Area of Science:
- Natural Language Processing
- Speech Recognition Technology
- Computational Linguistics
Background:
- End-to-end Chinese speech recognition models often output characters directly, leading to lower performance compared to other languages.
- The complexity of the Chinese language, particularly the intricate relationship between text and pronunciation, poses a significant challenge.
- Existing models struggle to effectively capture the nuances of Chinese phonetics, impacting overall accuracy.
Purpose of the Study:
- To enhance Chinese speech recognition accuracy by addressing the complexities of character-pronunciation relationships.
- To introduce a novel model inspired by the phonetic learning process of human beginners.
- To reduce character error rates in Chinese speech recognition systems.
Main Methods:
- Proposed the Syllable-Character Collaborative Model (SCCM), integrating phonetic elements like initials, finals, and pinyin into the training process.
- Developed a Pinyin-Ensemble module utilizing ensemble learning to minimize pinyin recognition errors.
- Incorporated phonetic information as auxiliary input, mimicking the learning progression of Chinese language novices.
Main Results:
- The SCCM demonstrated superior performance compared to prior end-to-end methods that used pinyin as auxiliary information.
- Achieved a significant 45.7% relative reduction in Character Error Rate (CER) compared to the AISHELL-1 baseline.
- Successfully reduced both pinyin and character error rates, indicating improved phonetic and textual recognition.
Conclusions:
- The Syllable-Character Collaborative Model (SCCM) effectively addresses the challenges of Chinese speech recognition by leveraging phonetic components.
- Integrating pinyin and syllable information, along with ensemble learning for pinyin, leads to substantial improvements in accuracy.
- The proposed approach offers a promising direction for advancing Chinese speech recognition technology.

