Related Experiment Video
Updated: Jul 19, 2026

Holistic Facial Composite Creation and Subsequent Video Line-up Eyewitness Identification Paradigm
Published on: December 24, 2015
VPT: Video portraits transformer for realistic talking face generation
Zhijun Zhang1, Jian Zhang2, Weijian Mai3
1School of Automation Science and Engineering, South China University of Technology, China; Key Library of Autonomous Systems and Network Control, Ministry of Education, China; Jiangxi Thousand Talents Plan, Nanchang University, China; College of Computer Science and Engineering, Jishou University, China; Guangdong Artificial Intelligence and Digital Economy Laboratory (Pazhou Lab), China; Shaanxi Provincial Key Laboratory of Industrial Automation, School of Mechanical Engineering, Shaanxi University of Technology, Hanzhong, China; School of Information Science and Engineering, Changsha Normal University, Changsha, China; School of Automation Science and Engineering, and also with the Institute of Artificial Intelligence and Automation, Guangdong University of Petrochemical Technology, Maoming, China; Key Laboratory of Large-Model Embodied-Intelligent Humanoid Robot (2024KSYS004), China; The Institute for Super Robotics (Huangpu), Guangzhou,, China.
This study introduces a new talking face generation method, Video Portraits Transformer (VPT), for realistic videos with identity preservation and natural blinks. The framework improves audiovisual synchronization and facial details for applications like digital assistants.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Existing audio-driven talking face generation methods struggle with photo-realism, identity preservation, and natural facial details like blinks.
- Synchronization between audio and video is a key challenge in current talking face synthesis.
Purpose of the Study:
- To propose a novel talking face generation framework, Video Portraits Transformer (VPT), addressing limitations in realism, identity preservation, and blink synchronization.
- To enhance the synthesis of photo-realistic talking face videos with controllable and natural blink movements.
Main Methods:
- The proposed Video Portraits Transformer (VPT) framework employs a two-stage process: audio-to-landmark and landmark-to-face.
- The audio-to-landmark stage utilizes a transformer encoder to predict facial landmarks from audio and Eye Aspect Ratio (EAR).
- The landmark-to-face stage uses a video-to-video (vid-to-vid) network for landmark-to-realistic video synthesis, incorporating a spontaneous blink generation module.
Main Results:
- The VPT method successfully generates photo-realistic talking face videos with high identity preservation and accurate audiovisual synchronization.
- The spontaneous blink generation module effectively mimics real blink duration distribution and frequency, adding naturalness to the generated videos.
- Extensive experiments validate the framework's ability to produce high-quality talking face videos with natural blink movements.
Conclusions:
- The Video Portraits Transformer (VPT) framework offers a significant advancement in audio-driven talking face generation, achieving superior realism and identity preservation.
- The integration of controllable blink movements enhances the naturalness and expressiveness of synthesized talking faces.
- This research contributes to more immersive and realistic digital interactions in applications like virtual assistants and video conferencing.
More Related Videos
06:53Creating Virtual-hand and Virtual-face Illusions to Investigate Self-representation
Published on: March 1, 2017
06:20Author Spotlight: Development of an Automated Camera-Based System for Real-Time Blast Overpressure Monitoring and TBI Risk Assessment in Military Training
Published on: December 6, 2024
Related Concept Videos
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Muscles for Facial Expressions
Transformers with Off-Nominal Turns Ratios
Facial Feedback Hypothesis
Transformation