Related Experiment Videos
Evaluating the real-world robustness of face-swap detection models under compression and noise
Li Baocai1, Fazli Bin Azzali1, Nik Fatinah Binti N Mohd Farid1
1School of Computing, Universiti Utara Malaysia, Sintok, Malaysia.
Introduction:
Recent advances in generative adversarial networks (GANs) and autoencoding techniques have significantly improved the realism of face-swap and deepfake media, creating substantial challenges for digital media authentication. Although existing deepfake detection models achieve high accuracy on benchmark datasets, their robustness under real-world media degradations remains insufficiently explored.
Methods:
This study systematically evaluates the resilience of five leading face-swap detection models-XceptionNet, MesoNet, FSD-GAN, FakeTracer, and a Hybrid + Landmark approach-under four common distortions: JPEG compression (quality levels 20-90), Gaussian noise (σ = 0.01-0.05), motion blur (kernel size 3-15), and video encoding artefacts (bitrate 50-500 kbps). Experiments were conducted using the FaceForensics++ dataset (1,000 videos: 720 training, 140 validation, and 140 testing) and Celeb-DF v2 (590 videos: 400 training, 90 validation, and 100 testing). Performance was assessed using accuracy, F1-score, area under the curve (AUC), and degradation rate (Δ) between clean and distorted conditions.
Results:
The results demonstrate a substantial reduction in detection performance under degraded conditions. Average accuracy declined from 94.7% on clean data to 67.8% on distorted data, corresponding to an overall degradation rate of -26.9%. JPEG compression and motion blur caused the most significant performance drops, with reductions of up to 35%, particularly for lightweight CNN-based detectors. In contrast, FSD-GAN and FakeTracer exhibited greater robustness, maintaining degradation rates of no more than -15% due to their latent fingerprinting and trace embedding mechanisms.
Discussion:
The findings highlight the limitations of current deepfake detection systems when deployed in real-world environments where media distortions are prevalent. The study emphasizes the need for distortion-aware training strategies, cross-condition benchmarking, and deployment-oriented evaluation protocols. Furthermore, a dual-branch framework integrating a Vision Transformer (ViT) for spatial artefact detection with a Recurrent Neural Network (RNN) or Temporal Convolutional Network (TCN) for temporal coherence modelling is proposed as a promising direction for improving the robustness and reliability of future deepfake detection systems.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...