Related Experiment Video
Updated: Jan 7, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Spoof detection with dynamic learnable sparse attention and tri-modal fusion in resource-constrained audio systems
Xinwei Wang1, Zhicheng Tan2, Guo Li2
1Department of Forensic Science, Fu Jian Police College, Fuzhou, Fujian, China.
This study introduces a new Dynamic Learnable Sparse Attention (DLSA) framework for efficient audio spoof detection in automatic speaker verification (ASV) systems. The DLSA method significantly reduces computational costs while improving detection accuracy on resource-constrained devices.
Area of Science:
- Speech Processing
- Machine Learning
- Signal Processing
Background:
- Automatic speaker verification (ASV) systems are vulnerable to audio spoofing attacks.
- Existing spoof detection methods, like Multi-Head Attention (MHA), are computationally expensive and memory-intensive, limiting their use in resource-constrained environments.
Purpose of the Study:
- To develop an efficient and robust spoof detection framework for resource-constrained ASV systems.
- To reduce the computational complexity and memory footprint of spoof detection methods.
Main Methods:
- Proposed a novel Dynamic Learnable Sparse Attention (DLSA) framework integrating Mel-Frequency Cepstral Coefficients (MFCC), Constant-Q Transform (CQT), and raw waveform.
- Employed a ResNet backbone for raw waveform feature extraction and a hybrid loss function (cross-entropy and center loss) for optimized feature representation.
- Implemented a learnable attention mechanism for dynamic cross-modal fusion of spectral and temporal features.
Main Results:
- Achieved an 80% reduction in computational costs compared to MHA-based methods.
- Attained an Equal Error Rate (EER) of 0.68% and a minimum tandem Detection Cost Function (t-DCF) of 0.0173 on the ASVspoof 2019 LA dataset.
- Demonstrated a 33.6% improvement in EER reduction over existing methods.
Conclusions:
- The DLSA framework offers an efficient and effective solution for spoof detection in ASV systems with limited resources.
- The proposed method enhances robustness against sophisticated audio spoofing techniques.
- This work paves the way for deploying advanced security features on edge devices.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Auditory Perception
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Nonconscious Mimicry
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
