Related Experiment Video
Updated: Dec 25, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.9K
End-to-End Automatic Pronunciation Error Detection Based on Improved Hybrid CTC/Attention Architecture
Long Zhang1, Ziping Zhao1, Chunmei Ma1
1College of Computer and Information Engineering, Tianjin Normal University, Tianjin 300387, China.
Sensors (Basel, Switzerland)
|March 29, 2020
Summary
This study introduces an advanced automatic pronunciation error detection (APED) system using a hybrid deep learning approach. The new method simplifies pronunciation training and achieves strong performance in identifying errors.
Area of Science:
- Speech Technology
- Artificial Intelligence
- Computational Linguistics
Background:
- Traditional automatic pronunciation error detection (APED) relies on complex automatic speech recognition (ASR) systems.
- Advancements in deep learning have led to mature end-to-end ASR technologies, offering new possibilities for APED.
- Existing APED methods often require force alignment and segmentation, increasing complexity.
Purpose of the Study:
- To develop a novel, simplified APED algorithm leveraging end-to-end ASR.
- To improve the performance and applicability of APED for computer-assisted pronunciation training (CAPT).
- To create an L1-independent solution for pronunciation error detection.
Main Methods:
- Constructed an end-to-end ASR system using a hybrid Connectionist Temporal Classification (CTC) and attention (seq2seq) architecture.
- Incorporated an adaptive parameter to enhance the synergy between CTC and attention models.
- Applied the improved ASR system to the Mandarin APED task.
Main Results:
- The hybrid CTC/attention ASR system achieved good results in Mandarin APED.
- The proposed APED method eliminates the need for force alignment, segmentation, and multiple complex models (acoustic, language).
- The system demonstrates accuracy comparable to state-of-the-art DNN-DNN ASR systems and superior F-measure performance for APED.
Conclusions:
- The developed end-to-end ASR-based APED system offers a convenient and straightforward solution.
- This approach is suitable for L1-independent computer-assisted pronunciation training (CAPT).
- The hybrid CTC/attention architecture provides a strong foundation for effective pronunciation error detection.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
13.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
13.9K
Improving Translational Accuracy
3.5K
3.5K

