Related Experiment Video
Updated: Jan 11, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
iWAX: interpretable Wav2vec-AASIST-XGBoost framework for voice spoofing detection
Seungeun Lee1,2, Sunmook Choi1,3, Taein Kang4
1Department of Mathematics, Korea University, Seoul, 02841, South Korea.
Abstract:
Recent advances in deep learning have led to the widespread use of pre-trained large-scale speech models, such as wav2vec 2.0 (w2v2), in voice spoofing detection. However, the interpretability of such models remains a critical challenge due to their complex internal representations. In this paper, we propose iWAX, an interpretable voice spoofing countermeasure that combines a fine-tuned w2v2 front-end with the AASIST back-end, and an XGBoost classifier. iWAX leverages the feature importance mechanism of XGBoost to identify which temporal segments and frequency bands of the audio w2v2 prioritizes during spoofing detection. To enable frequency-based interpretability, we apply sinc filters to isolate specific spectral regions of input raw waveforms. Temporal analysis is conducted by selecting key features extracted from w2v2 and analyzing their contribution across time. Experimental results on the ASVspoof 2019 LA dataset demonstrate that iWAX not only outperforms baseline models such as AASIST and w2v2-AASIST but also provides human-understandable explanations of its predictions. Further analysis with LightGBM validates the robustness of our approach across different boosting models. Overall, iWAX offers a compelling balance between interpretability and performance, addressing the limitations of both traditional machine learning and modern deep learning-based countermeasures.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Impression Management Techniques IV: Altercasting
Air-entraining Agents
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...

