Related Experiment Video
Updated: Oct 11, 2026

Measuring Connectivity in the Primary Visual Pathway in Human Albinism Using Diffusion Tensor Imaging and Tractography
Published on: August 11, 2016
SDMatte++: Grafting High-Resolution Diffusion Models for Interactive Matting
Abstract:
Recent advances in interactive image matting have achieved substantial progress in capturing the primary regions of foreground objects. However, limited generalization capability and inadequate preservation of fine-grained details remain major challenges in interactive image matting. Meanwhile, diffusion models trained on large-scale image-text pairs learn unified visual representations, which enable them to provide strong semantic priors, making them an attractive solution for interactive image matting. To this end, we propose SDMatte, a diffusion-driven interactive matting framework that leverages the strong semantic priors of diffusion models to alleviate the limited generalization capability of existing interactive image matting methods. Specifically, we transform the text-driven interaction capability of diffusion models into a visual prompt-driven interaction capability to enable interactive matting. Furthermore, we integrate coordinate embeddings and opacity embeddings of target objects into U-Net, enhancing the model's sensitivity to spatial position information and opacity information. In addition, we propose a masked self-attention mechanism that enables the model to focus on regions specified by visual prompts, leading to improved interactive performance. Nevertheless, despite the advantages of SDMatte, its large-scale architecture typically necessitates training at low resolutions, which limits its ability to effectively leverage the richer and more pronounced high-frequency information available in high-resolution images for modeling fine-grained details. To this end, we propose a high-resolution detail injection strategy that progressively incorporates high-frequency information extracted from higher-resolution images into low-resolution latent features during the upsampling process, thereby enhancing the model's capability to preserve fine-grained details. Finally, due to the substantial computational overhead of diffusion-based large-scale architectures, which poses challenges for practical deployment, we distill SDMatte and SDMatte++ to obtain lightweight variants, LiteSDMatte and LiteSDMatte++. With only slight performance degradation, the lightweight variants achieve approximately 38% reduction in parameters, around 80% reduction in FLOPs, and roughly 3× faster inference compared to the original models. Extensive experiments on multiple datasets demonstrate the superior performance of our method, validating its generalization capability, fine-grained detail preservation, and accuracy in interactive matting.

