Related Experiment Video
Updated: Jan 11, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
iWAX: interpretable Wav2vec-AASIST-XGBoost framework for voice spoofing detection.
Seungeun Lee1,2, Sunmook Choi1,3, Taein Kang4
1Department of Mathematics, Korea University, Seoul, 02841, South Korea.
This study introduces iWAX, an interpretable system for voice spoofing detection. It uses a deep learning model (wav2vec 2.0) and XGBoost to explain its decisions, outperforming existing methods.
Area of Science:
- Speech processing
- Machine learning
- Artificial intelligence
Background:
- Deep learning models like wav2vec 2.0 (w2v2) are increasingly used for voice spoofing detection.
- However, their complex nature hinders interpretability, making it difficult to understand their decision-making process.
Purpose of the Study:
- To develop an interpretable voice spoofing countermeasure (iWAX) that combines w2v2 with XGBoost.
- To enable explanations of detection predictions by identifying important temporal and frequency segments.
Main Methods:
- iWAX utilizes a fine-tuned w2v2 front-end and AASIST back-end with an XGBoost classifier.
- Sinc filters are applied for frequency band analysis, and temporal analysis focuses on key w2v2 features.
- XGBoost's feature importance mechanism is central to the interpretability approach.
Main Results:
- iWAX demonstrated superior performance compared to baseline models (AASIST, w2v2-AASIST) on the ASVspoof 2019 LA dataset.
- The system provided human-understandable explanations for its voice spoofing detection predictions.
- Robustness was confirmed using LightGBM, indicating broad applicability of the interpretability method.
Conclusions:
- iWAX achieves a strong balance between high performance and interpretability in voice spoofing detection.
- The approach addresses the limitations of traditional and deep learning-based countermeasures.
- This work facilitates trust and understanding in AI-driven audio security systems.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Impression Management Techniques IV: Altercasting
Air-entraining Agents
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...

