Related Experiment Video
Updated: Sep 13, 2025

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
471
Phoneme-Aware Hierarchical Augmentation and Semantic-Aware SpecAugment for Low-Resource Cantonese Speech Recognition
Lusheng Zhang1,2, Shie Wu1,2, Zhongxun Wang1,2
1School of Physics and Electronic Information, Yantai University, Yantai 264005, China.
Sensors (Basel, Switzerland)
|July 30, 2025
Summary
This study introduces a phoneme-aware framework to improve Cantonese Automatic Speech Recognition (ASR) without extra data. The method enhances pronunciation diversity and contextual understanding, significantly reducing character error rates for low-resource tonal languages.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Computational Linguistics
Background:
- Cantonese Automatic Speech Recognition (ASR) faces challenges due to tonal complexity, acoustic variations, and limited labeled data.
- Existing ASR models struggle with the nuances of tonal languages, impacting performance in real-world applications.
- The need for robust ASR systems for low-resource languages is critical for broader accessibility and technological integration.
Purpose of the Study:
- To develop a phoneme-aware hierarchical augmentation framework to enhance Cantonese ASR performance.
- To improve ASR accuracy without requiring additional manual annotations or transcriptions.
- To create a model-independent approach applicable to other low-resource tonal languages.
Main Methods:
- Implemented a Phoneme Substitution Matrix (PSM) using Montreal Forced Aligner and Tacotron-2 to introduce phonetic variations.
- Employed a semantic-aware SpecAugment strategy leveraging wav2vec 2.0 attention and keyword boundaries for adaptive masking.
- Utilized a reinforcement-learning controller for online tuning of the masking schedule to encourage broader contextual reliance.
Main Results:
- Reduced character error rate (CER) on the Common Voice Cantonese 50 h subset from 26.17% to 16.88% (wav2vec 2.0) and 38.83% to 23.55% (Zipformer).
- Achieved further CER reductions at 100 h, reaching 4.27% and 2.32% respectively, demonstrating significant performance gains (32-44% relative).
- Ablation studies confirmed the complementary benefits of both phoneme-level augmentation and adaptive masking techniques.
Conclusions:
- The proposed framework offers a practical and effective solution for improving Cantonese ASR accuracy.
- The phoneme-aware hierarchical augmentation is model-independent, providing a versatile approach for low-resource tonal languages.
- The intelligent sensing-oriented framework shows potential for deployment in edge devices and voice-interactive systems.

