Related Experiment Video
Updated: Aug 6, 2026

09:49
Holistic Facial Composite Creation and Subsequent Video Line-up Eyewitness Identification Paradigm
Published on: December 24, 2015
Multi-condition guided diffusion model for face sketch-to-photo synthesis
Yue Que1, Xuegui Cheng1, Shuqian Shi1
1School of Information and Software Engineering, East China Jiaotong University, Nanchang, 330013, China.
Summary
This study introduces a novel diffusion-based framework for facial sketch synthesis, improving structural accuracy and identity preservation. The method enhances realism and detail reconstruction, outperforming existing generative adversarial networks.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Digital Forensics
Background:
- Facial sketch synthesis is crucial for cross-modal face analysis and digital forensics.
- Existing generative adversarial network (GAN)-based methods struggle with training instability, structural distortions, and identity inconsistency, especially with limited paired data.
Purpose of the Study:
- To propose a diffusion-based framework to enhance structural and textural fidelity in facial sketch synthesis.
- To address limitations of existing methods regarding training stability and detail reconstruction.
Main Methods:
- A stage-wise multi-condition guidance mechanism is employed within a diffusion framework.
- Semantic segmentation features guide structural information during downsampling.
- A hybrid cross-attention mechanism integrates textures and denoised noise during upsampling.
- A Vision Transformer is integrated into the U-Net backbone for improved global contextual information capture.
Main Results:
- The proposed method demonstrates strong performance in SSIM and FSIM, indicating improved structural similarity and feature-level fidelity.
- Competitive LPIPS results show enhanced perceptual similarity.
- The framework achieves competitive visual coherence and identity preservation compared to recent baselines.
Conclusions:
- The diffusion-based framework effectively enhances structural and textural fidelity in facial sketch synthesis.
- The stage-wise guidance and Vision Transformer integration contribute to improved realism and detail reconstruction.
- This approach offers a promising solution for generating high-fidelity facial sketches, particularly under data constraints.
