Related Experiment Video
Updated: Jun 28, 2025

Creating Virtual-hand and Virtual-face Illusions to Investigate Self-representation
Published on: March 1, 2017
Face anti-spoofing with cross-stage relation enhancement and spoof material perception
Daiyuan Li1, Guo Chen2, Xixian Wu3
1South China University of Technology, Guangzhou, 510006, Guangdong, China; Pazhou Laboratory, Guangzhou, 510000, Guangdong, China; Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, Guangzhou, 510000, Guangdong, China.
Abstract:
Face Anti-Spoofing (FAS) seeks to protect face recognition systems from spoofing attacks, which is applied extensively in scenarios such as access control, electronic payment, and security surveillance systems. Face anti-spoofing requires the integration of local details and global semantic information. Existing CNN-based methods rely on small stride or image patch-based feature extraction structures, which struggle to capture spatial and cross-layer feature correlations effectively. Meanwhile, Transformer-based methods have limitations in extracting discriminative detailed features. To address the aforementioned issues, we introduce a multi-stage CNN-Transformer-based framework, which extracts local features through the convolutional layer and long-distance feature relationships via self-attention. Based on this, we proposed a cross-attention multi-stage feature fusion, employing semantically high-stage features to query task-relevant features in low-stage features for further cross-stage feature fusion. To enhance the discrimination of local features for subtle differences, we design pixel-wise material classification supervision and add a auxiliary branch in the intermediate layers of the model. Moreover, to address the limitations of a single acquisition environment and scarcity of acquisition devices in the existing Near-Infrared dataset, we create a large-scale Near-Infrared Face Anti-Spoofing dataset with 380k pictures of 1040 identities. The proposed method could achieve the state-of-the-art in OULU-NPU and our proposed Near-Infrared dataset at just 1.3GFlops and 3.2M parameter numbers, which demonstrate the effective of the proposed method.
More Related Videos
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Nonconscious Mimicry
Self-Evaluation: Self-Enhancement and Self-Verification
Self-Presentation: Self-Monitoring and Self-Handicapping
Stereotype Content Model