Related Experiment Video
Updated: Sep 17, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
LSTM autoencoder based parallel architecture for deepfake audio detection with dynamic residual encoding and feature
Priyanka Muruganandham1, Govardhana Rajan Thangasamy1, Sangeetha Jayaraman2
1Department of CSE, Srinivasa Ramanujan Centre, SASTRA Deemed to be University, Kumbakonam, 612 001, India.
This study introduces a new deepfake audio detection model, LSTM-AE-DRDE, improving accuracy by analyzing temporal cues and audio features. The advanced framework effectively distinguishes real from fake audio, enhancing security against misinformation.
Area of Science:
- Computer Science
- Artificial Intelligence
- Signal Processing
Background:
- Deepfake audio detection is crucial due to advancements in synthetic speech technology.
- Existing models face challenges like temporal inconsistencies and weak contextual representation.
Purpose of the Study:
- To enhance deepfake audio detection performance.
- To address limitations in current detection models.
Main Methods:
- Proposed a novel framework: Long Short-Term Memory Auto-Encoder with Dynamic Residual Difference Encoding (LSTM-AE-DRDE).
- Utilized attention-enhanced LSTM with contrastive learning for temporal cues.
- Employed residual differences across audio variants for real-vs-fake separability.
- Integrated diverse speech features: MFCC, temporal, prosodic, wavelet packet, and glottal parameters.
Main Results:
- Achieved high classification accuracies on five benchmark datasets (90%-97%).
- Demonstrated superior performance compared to eleven state-of-the-art methods with an ROC-AUC of ~98%.
- Feature-wise ablation study confirmed the framework's robustness.
Conclusions:
- The proposed LSTM-AE-DRDE framework significantly improves deepfake audio detection.
- Integration of diverse features and novel encoding methods enhances model reliability and performance.
Related Concept Videos
Parallel Resonance
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Elaborative Rehearsals
The effectiveness of...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
