Related Experiment Video
Updated: Jun 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improved Speech Authenticity Detection in Chinese-English Bilingual Contexts
Cheng-Yuan Tsai1, Sheng-Chain Chang2, Chao-Hsiang Hung2
1Forensic Science Division, Ministry of Justice Investigation Bureau, New Taipei City 231, Taiwan.
This study introduces a novel deep learning model for detecting tampered audio, significantly improving accuracy in multilingual environments. The enhanced ResNet-LSTM approach surpasses current leading models in identifying sophisticated voice spoofing attacks.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Cybersecurity
Background:
- The proliferation of voice technology necessitates advanced methods for authenticating speech.
- Existing audio tampering detection systems often lack robustness across different languages and tampering types.
- Spoofing attacks pose a significant threat to voice-based security systems.
Purpose of the Study:
- To develop and evaluate an improved model for detecting tampered audio, specifically addressing challenges in multilingual settings.
- To enhance the generalization capabilities of deep learning models for audio tampering detection.
- To benchmark the proposed model against state-of-the-art methods in recent competitions.
Main Methods:
- Developed a hybrid deep learning model integrating an enhanced ResNet architecture with a Long Short-Term Memory (LSTM) network.
- Created a bilingual dataset combining self-recorded Chinese speech and public English audio samples (VCTK2).
- Evaluated the model using advanced tampering techniques like CycleGAN voice conversion and auto splicing.
Main Results:
- The proposed ResNet-LSTM model achieved a superior performance with an equal error rate (EER) of 11.62% on a bilingual dataset.
- The model demonstrated superior performance compared to leading approaches from the ASVSpoof 2021 and ADD 2022 competitions.
- Effectiveness was validated against realistic tampering scenarios, including voice conversion and audio splicing.
Conclusions:
- The integrated ResNet-LSTM model offers a robust solution for detecting tampered audio, particularly in challenging multilingual contexts.
- The approach significantly advances the state-of-the-art in anti-spoofing technology.
- This work provides a foundation for more secure voice communication systems.
More Related Videos
08:32Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Improving Translational Accuracy
Language and Cognition