Related Experiment Videos
MODdapter: Spatially-aware text embeddings for zero-shot semantic segmentation
Jiaxiang Fang1, Shiqiang Ma2, Jing Wang3
1School of Computer Science and Engineering, Central South University, Changsha, 410017, China; Sony Research and Development Center Beijing Lab, Chao-Yang District, Beijing, 100027, China.
Abstract:
Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegmentation. To address this, we propose MODdapter, a Multimodal Adapter with Proportional-Integral-Derivative (PID) control for precise dual-space alignment of vision and language features. Our core idea is to leverage visual cues obtained during testing to complement textual information, improving the alignment between text and visual features. MODdapter extracts visual semantic cues from coarse segmentation results and integrates them as supplementary textual data, allowing for a more accurate projection of text features into the feature space and enhancing fine-grained recognition. To maintain stability and counteract potential oscillations caused by deviations in the extracted instance cues, we incorporate a PID control module that regulates the alignment process. PID control module leverages the combined joint constraints of short-term (derivative) and long-term (integral) memory, achieving fine-grained recognition of unseen target. This strategy significantly enhances the recognition of unseen objects, achieving state-of-the-art performance across multiple datasets. On the challenging COCO-20i dataset, our method outperforms the current state-of-the-art by a significant improvements of 6.9%.