Related Experiment Video
Updated: Jun 4, 2026

Creating Virtual-hand and Virtual-face Illusions to Investigate Self-representation
Published on: March 1, 2017
MaPPA: Multimodal Controllable Person Image Generation With Pose and Appearance Guidance
Abstract:
Person image generation has become an increasingly important problem in computer vision with broad applications in virtual try-on, digital content creation, entertainment, and human-computer interaction. Despite recent advances in diffusion-based generative models that can produce photorealistic results, existing pipelines are still constrained by rigid modality requirements. Most prior methods rely on fixed patterns such as pose-plus-appearance inputs or text-only descriptions, limiting their flexibility and controllability in practical scenarios where users may prefer different or mixed types of guidance. This lack of adaptability poses a clear barrier to deployment in real-world systems. To overcome these challenges, we propose MaPPA (Multimodal Controllable Person Image Generation with Pose and Appearance Guidance), a unified multimodal framework that leverages transformer-based latent diffusion models. The central idea is to provide composable control over multiple modalities, enabling person image generation conditioned on text prompts, reference appearance images, pose keypoints, or any combination thereof. To achieve this, we introduce a unified framework that incorporates dedicated control blocks for appearance and pose guidance, which are strategically interleaved with transformer base blocks. These control blocks are conditioned on a unified multimodal embedding that integrates heterogeneous inputs into a consistent representation, thereby supporting arbitrary modality combinations with a single unified pipeline. Another key contribution of MaPPA is a cumulative classifier-free guidance strategy that enables allowing users to independently adjust the strength of appearance and pose guidance via scalar weights from different control signals. This design allows users to adjust the relative strength of appearance versus pose guidance, providing fine-grained controllability during inference. Furthermore, to address the common problem of detail loss in latent-diffusion decoders, we propose a texture enhancement decoding (TED) strategy, which fine-tunes the VAE decoder with edge-aware reconstruction objectives. This refinement significantly alleviates texture distortion, preserving high-frequency details in clothing patterns, facial regions, and other fine structures. Extensive experiments confirm that MaPPA achieves competitive quantitative scores while providing superior perceptual quality and user preference. Unlike task-specific methods, our framework-though constrained by the capacity of the unified multimodal embedding-supports combinations of text, pose, and appearance modalities within a single pipeline, demonstrating both flexibility and practical value. The code and the corresponding model will be made publicly available to the research community.
Related Concept Videos
Impression Management Techniques I: Managing Appearances
Modeling and Similitude