Related Experiment Video
Updated: Jan 7, 2026

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
5.2K
Audio Deepfake Detection via a Fuzzy Dual-Path Time-Frequency Attention Network
Jinzi Li1, Hexu Wang1,2, Fei Xie3,4
1Xi'an Key Laboratory of Human-Machine Integration and Control Technology for Intelligent Rehabilitation, Xijing University, Xi'an 710123, China.
Sensors (Basel, Switzerland)
|December 31, 2025
Summary
This study introduces a novel Dual-Path Time-Frequency Attention Network (DPTFAN) to combat audio deepfakes. The method enhances detection accuracy and robustness against noise and compression, improving information security.
Area of Science:
- Artificial Intelligence
- Information Security
- Signal Processing
Background:
- Audio deepfake technology poses significant threats to information security due to advancements in speech synthesis.
- Existing detection methods struggle with robustness against environmental noise, signal compression, and subtle fake audio features.
- Highly concealed fake audio remains difficult to identify effectively with current techniques.
Purpose of the Study:
- To develop a robust and accurate method for detecting audio deepfakes.
- To address the limitations of existing methods in handling noisy and compressed audio.
- To improve the identification of highly concealed fake audio through advanced feature characterization and enhancement.
Main Methods:
- Proposes a Dual-Path Time-Frequency Attention Network (DPTFAN) incorporating Pythagorean Hesitant Fuzzy Sets (PHFS) for uncertainty modeling.
- Employs a dual-path attention mechanism in time and frequency domains to improve feature representation and discriminative power.
- Introduces a Lightweight Fuzzy Branch Network (LFBN) for explicit enhancement of ambiguous audio features, balancing performance and efficiency.
Main Results:
- Achieved 98.94% accuracy on the ASVspoof 2019 LA dataset.
- Reached 99.40% accuracy on the FoR (Fake or Real) dataset.
- Demonstrated superior performance and robustness compared to existing mainstream audio deepfake detection methods.
Conclusions:
- The proposed DPTFAN method offers excellent detection performance and robustness for audio deepfakes.
- Uncertainty modeling with PHFS and dual-path attention effectively characterizes and enhances fake audio features.
- The LFBN contributes to improved performance while maintaining computational efficiency in deepfake detection.
