Related Experiment Video
Updated: Jul 8, 2026

07:34
Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
17.3K
Enhanced Multi-Scale Cross-Attention for Person Image Generation
Summary
This study introduces XingGAN, a novel generative adversarial network (GAN) for person image generation. XingGAN utilizes cross-attention mechanisms to improve appearance and shape synthesis, achieving faster training and inference than diffusion models.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Deep Learning
Background:
- Generative Adversarial Networks (GANs) are widely used for image generation.
- Person image generation presents challenges in synthesizing realistic appearance and shape.
- Existing GANs often struggle with effectively integrating multi-modal features for complex generation tasks.
Purpose of the Study:
- To propose a novel cross-attention-based GAN, named XingGAN, for improved person image generation.
- To enhance the fusion of appearance and shape information for more accurate synthesis.
- To develop a computationally efficient method that rivals diffusion-based model performance.
Main Methods:
- Developed XingGAN with two generation branches for appearance and shape.
- Introduced novel cross-attention blocks for feature transfer and embedding updates.
- Implemented multi-scale cross-attention blocks for long-range correlation learning.
- Proposed an enhanced attention (EA) module to refine attention weights.
- Integrated a densely connected co-attention module for feature fusion.
Main Results:
- XingGAN outperforms existing GAN-based methods in person image generation.
- The proposed method achieves performance comparable to diffusion-based models.
- XingGAN demonstrates significantly faster training and inference speeds compared to diffusion models.
Conclusions:
- The novel cross-attention mechanisms in XingGAN effectively improve person image synthesis.
- XingGAN offers a compelling alternative to diffusion models, balancing performance and efficiency.
- This work advances GAN-based approaches for complex image generation tasks.

