Related Experiment Video
Updated: Apr 22, 2026

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
1.3K
Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE
Prateek Agrawal1,2, Dharmendra Pathak3, Vishu Madaan1,4
1School of Computer Science and Engineering, Rungta International Skills University, Bhilai, Chhattisgarh, 490024, India.
Scientific Reports
|April 20, 2026
Summary
This study benchmarks deep learning models for deepfake video detection. 3D Convolutional Neural Network (3DCNN) achieved the highest accuracy, but generalization remains a challenge for all models.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Digital Forensics
Background:
- Advanced AI techniques, including Generative Adversarial Networks (GANs), create increasingly realistic deepfake videos, posing significant threats to digital security, misinformation dissemination, and privacy.
- The sophistication of deepfakes makes them difficult to distinguish from genuine media, necessitating robust detection and classification methods.
- Easy access to deepfake tools exacerbates risks in political manipulation, financial deception, and identity theft.
Purpose of the Study:
- To systematically benchmark and evaluate the generalization performance of four deep learning models for deepfake video detection and classification.
- To investigate the effectiveness of 3D Convolutional Neural Network (3DCNN), 3D Residual Network (3DResNet), Temporal Convolutional Network (TCN), and Variational Autoencoder (VAE) in identifying manipulated videos.
Main Methods:
- Trained and evaluated 3DCNN, 3DResNet, TCN, and VAE models on diverse datasets: FaceForensics++ (FF++), Deepfake Detection Challenge (DFDC), and Celeb-DF (CDF), comprising over 3500 real and fake video samples.
- Conducted experiments on an NVIDIA DGX A100 workstation for efficient model training.
- Focused on benchmarking and generalization study rather than proposing novel architectures.
Main Results:
- 3DCNN achieved the highest test accuracy of 64.68% under cross-dataset conditions, outperforming 3DResNet, TCN, and VAE.
- Significant degradation in model generalization and varied failure behaviors were observed when models encountered heterogeneous data distributions.
- The study highlights limitations in the generalization capabilities of commonly used spatiotemporal models for deepfake detection.
Conclusions:
- While 3DCNN shows promise, current deep learning models struggle with robust generalization across different datasets.
- The findings underscore the need for more dependable deepfake detection systems capable of handling diverse and evolving forgery techniques.
- This research provides critical insights for developing practical, real-world deepfake detection solutions.