Related Experiment Video
Updated: Jun 29, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
530
Transformer-based cascade networks with spatial and channel reconstruction convolution for deepfake detection.
Xue Li1, Huibo Zhou1, Ming Zhao1
1School of Mathematical Sciences, Harbin Normal University, Harbin 150025, China.
Mathematical Biosciences and Engineering : MBE
|March 29, 2024
Summary
This study introduces an advanced deepfake detection model using spatial and channel reconstruction convolution (SCConv) and vision transformers. The model achieves high accuracy in identifying fake videos, enhancing national security against forged media.
Area of Science:
- Computer Science
- Artificial Intelligence
- Cybersecurity
Background:
- The proliferation of sophisticated forged video technology poses significant threats to individuals, society, and national security.
- The rapid advancement of deepfake technology necessitates the development of robust and up-to-date detection models.
- Training deepfake detection models requires substantial volumes of diverse data, presenting a data management challenge.
Purpose of the Study:
- To propose a novel deepfake detection model that addresses the limitations of existing methods.
- To enhance the accuracy and efficiency of deepfake detection through innovative network architecture.
- To provide a reliable solution for identifying forged videos in the digital landscape.
Main Methods:
- A cascade network architecture integrating spatial and channel reconstruction convolution (SCConv) with a vision transformer was developed.
- The model utilizes SCConv and regular convolution in its initial stages for fake video detection, followed by a vision transformer.
- The feed-forward layer of the vision transformer was enhanced to improve detection accuracy and reduce computational load.
- Datasets were processed by splitting video frames and extracting faces to create comprehensive real and fake face image datasets.
Main Results:
- The proposed model achieved high detection accuracies across multiple benchmark datasets: 87.92% on DFDC, 99.23% on FaceForensics++, and 99.98% on Celeb-DF.
- The enhanced vision transformer component successfully increased detection accuracy while simultaneously reducing the model's computational requirements.
- Qualitative assessments, including visualization results, demonstrated the model's effectiveness in authenticating video content.
Conclusions:
- The developed deepfake detection model, leveraging SCConv and an enhanced vision transformer, offers a highly effective solution for identifying forged videos.
- The model's superior performance on diverse datasets and its computational efficiency make it a valuable tool in combating the spread of malicious deepfakes.
- The study validates the model's efficacy through rigorous testing and confirms its potential for real-world application in deepfake detection.

