Related Experiment Video
Updated: Jan 9, 2026

06:19
Integration of Animal Behavioral Assessment and Convolutional Neural Network to Study Wasabi-Alcohol Taste-Smell Interaction
Published on: August 16, 2024
795
STFANet: A spatial and temporal feature aggregation network for fake face detection in videos
Guoren Yao1, Gaoming Yang2, Xintian Liu3
1School of Computer Science, Huainan Normal University, Huainan, Anhui, China.
Plos One
|December 10, 2025
Summary
Detecting manipulated videos is challenging. The Spatial and Temporal Feature Aggregation Network (STFANet) effectively identifies facial forgeries by integrating spatial and temporal features, achieving high accuracy on benchmark datasets.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Digital Forensics
Background:
- Rapid advancements in video synthesis technologies make video authenticity verification increasingly difficult.
- Current detection methods often overlook comprehensive spatio-temporal characteristics, relying mainly on intra-frame spatial artifacts and temporal inconsistencies.
- This limitation restricts the ability to effectively detect sophisticated video manipulations.
Purpose of the Study:
- To propose a novel network, the Spatial and Temporal Feature Aggregation Network (STFANet), for enhanced video forgery detection.
- To address the limitations of existing methods by effectively exploiting spatio-temporal features.
- To improve the accuracy and robustness of manipulated video detection, particularly for facial forgeries.
Main Methods:
- Developed a two-path network structure within STFANet to independently extract spatial and temporal features.
- Integrated extracted spatial and temporal features to create high-fidelity spatio-temporal representations.
- Incorporated a Vision Transformer module to capture global dependencies within feature maps for enhanced representation.
Main Results:
- STFANet demonstrated high efficacy in detecting facial forgeries in videos.
- Achieved excellent performance on benchmark datasets: an AUC score of 0.9933 on FaceForensics++ and 0.9829 on Celeb-DF.
- Analysis confirmed that feature aggregation significantly improves the quality of spatio-temporal representations.
Conclusions:
- The proposed STFANet effectively addresses the challenges in video authenticity verification.
- The network's ability to aggregate spatial and temporal features, enhanced by a Vision Transformer, leads to superior forgery detection performance.
- STFANet represents a significant advancement in detecting manipulated videos, particularly facial forgeries.
Related Concept Videos
Association Areas of the Cortex
8.7K
Association areas are regions of the cerebral cortex that do not have a specific sensory or motor function. Instead, they integrate and interpret information from various sources to enable higher cognitive processes such as memory, learning, and decision-making. Some key association areas include the following:
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
8.7K
Masking and Demasking Agents
3.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
3.4K