Related Experiment Video
Updated: Jun 28, 2025

11:15
fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
7.2K
Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan1, Jun Xue1, Jianhua Tao2
1Anhui Province Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University, Hefei, 230601, China.
Summary
This study introduces a new method for fake speech detection (FSD) using fundamental frequency (F0) subbands. The novel approach significantly improves detection accuracy, achieving state-of-the-art results on a key benchmark dataset.
Area of Science:
- Speech Processing
- Artificial Intelligence
- Signal Analysis
Background:
- Synthetic speech often lacks the natural rhythm and fundamental frequency (F0) variations of genuine human speech.
- The fundamental frequency (F0) is a crucial feature for distinguishing between real and synthetic speech.
- Existing fake speech detection (FSD) methods may not fully leverage the discriminative potential of F0 variations.
Purpose of the Study:
- To propose a novel method for fake speech detection (FSD) by focusing on the fundamental frequency (F0) subband.
- To enhance the modeling of F0 subbands for improved FSD performance.
- To achieve state-of-the-art results in fake speech detection.
Main Methods:
- A novel F0 subband approach is proposed for fake speech detection (FSD).
- The spatial reconstructed local attention Res2Net (SR-LA Res2Net) architecture is introduced to effectively model F0 subbands.
- Res2Net backbone extracts multiscale information, enhanced by spatial reconstruction to preserve channel data and local attention to focus on F0 subband details.
Main Results:
- The proposed method achieved an equal error rate (EER) of 0.47% on the ASVspoof 2019 LA dataset.
- The minimum tandem detection cost function (min t-DCF) reached 0.0159.
- The system demonstrated state-of-the-art performance compared to other single systems.
Conclusions:
- The proposed F0 subband method, utilizing SR-LA Res2Net, is highly effective for fake speech detection (FSD).
- The integration of spatial reconstruction and local attention significantly enhances the modeling of F0 features.
- The achieved results represent a significant advancement in the field of fake speech detection.

