Related Experiment Videos
Mitigating Textual Noise in Multimodal FGVC via Hierarchical Semantic Purification and Multi-Stage Alignment
Abstract:
Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous textual descriptions: existing methods rely on raw or generated descriptions without filtering, introducing redundant and ambiguous semantic noise; and 2) Underutilization of hierarchical visual features: most approaches align single-layer visual features with auxiliary semantic embeddings, underutilizing hierarchical information. To address these challenges, we propose a task-oriented multimodal FGVC framework that eliminates textual redundancy while enhancing multi-layer alignment between cross-modalities. Specifically, our method comprises two key components: Hierarchical Semantic Purification (HSP) and Multi-layer Cross-Modal Alignment (MCA). The former employs a semantic distillation dictionary to eliminate redundant elements and uses a self-attention mechanism for ranking and semantic refinement. The latter establishes effective cross-modal fusion by integrating multi-layer features with purified text features, effectively combining multi-scale visual representations. Experimental results on 7 public datasets demonstrate that our proposed method outperforms existing counterparts, contributing to advancements in fine-grained visual classification.