MaPPA: Multimodal Controllable Person Image Generation With Pose and Appearance Guidance

Summary

This study introduces MaPPA, a multimodal framework for person image generation, enabling flexible control using text, pose, and appearance guidance. MaPPA offers enhanced controllability and detail preservation for realistic virtual try-on and content creation applications.