Related Experiment Video
Updated: Oct 24, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Non-native acoustic modeling for mispronunciation verification based on language adversarial representation learning.
Longfei Yang1, Kaiqi Fu2, Jinsong Zhang2
1Department of Information and Communication Engineering, Tokyo Institute of Technology, Tokyo, Japan.
This study introduces a novel pre-trained approach for non-native mispronunciation verification, leveraging native language speech data to overcome data sparsity in computer-aided pronunciation training (CAPT). The method effectively improves pronunciation error detection and feedback for language learners.
Area of Science:
- Speech Processing
- Computational Linguistics
- Language Acquisition
Background:
- Non-native mispronunciation verification is crucial for computer-aided pronunciation training (CAPT) systems.
- Existing methods face data sparsity issues due to the difficulty of collecting and annotating non-native speech data.
- This limits the effectiveness of current pronunciation feedback for language learners.
Purpose of the Study:
- To propose a pre-trained approach for non-native mispronunciation verification that utilizes speech data from both the learner's native and target languages.
- To address the data sparsity problem inherent in traditional CAPT systems.
- To enhance the accuracy and utility of pronunciation error detection and feedback.
Main Methods:
- An unsupervised model was developed to extract knowledge from large-scale unlabeled target language speech data.
- Language adversarial training was employed using the learner's native language to align feature distributions.
- A sinc filter was incorporated to capture formant-like features, aiding in articulation analysis.
Main Results:
- The pre-trained model effectively utilized knowledge from native language speech for non-native phone recognition and mispronunciation verification.
- Language adversarial representation learning significantly improved performance in these tasks.
- The inclusion of formant-like features via sinc filters further enhanced mispronunciation verification accuracy.
Conclusions:
- The proposed unsupervised pre-training approach effectively leverages native language speech data to improve non-native mispronunciation verification.
- Language adversarial training and formant-like feature extraction are key components for enhancing CAPT system performance.
- This method offers a promising solution for more accurate and instructive pronunciation feedback in language learning.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Impression Management Techniques IV: Altercasting

