Related Experiment Video
Updated: Sep 2, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Pose-Star++: Semantic-Visual Understanding for Fine-Grained Fashion Image Editing
Abstract:
Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are semantically flexible but structurally inaccurate. To bridge this gap, we propose Pose-Star++, a training-free, plug-and-play framework that introduces two core innovations: an LVLM-based Understanding Module that shifts from word- to sentence-level semantic-visual comprehension, eliminating cumbersome instruction pre-parsing and enabling robust understanding of complex natural language; a Bidirectional Calibration Module that co-optimizes semantic and structural constraints through forward pose-guided and backward attention-guided refinement, achieving precise, whole-body-reachable region calibration even under challenging in-the-wild poses. We further contribute the first real-world-oriented fashion-editing benchmark with diverse data, instructions, and tasks, exposing long-overlooked practical challenges. Extensive experiments demonstrate that Pose-Star++ significantly outperforms existing methods in semantic alignment, pose robustness, and in-the-wild generalization across complex scenarios, advancing toward practical, user-guided fashion creation.
Related Concept Videos
Modeling and Similitude
Stereotype Content Model
Vision
Visual System
Once through the pupil, the light passes through the lens, a...