Related Experiment Video
Updated: Aug 6, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
HDG-CLIP: Hierarchical Dual-Granularity Vision-Semantic Alignment for Open-Vocabulary Multi-Label Image
Summary
This study introduces Hierarchical Dual-Granularity Alignment-CLIP (HDG-CLIP) for Open-Vocabulary Multi-Label Image Classification (OV-MLIC). HDG-CLIP enhances recognition of unseen categories by addressing category coupling and scale variation, achieving state-of-the-art results.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Open-Vocabulary Multi-Label Image Classification (OV-MLIC) is crucial for real-world image recognition.
- Existing Vision and Language Pre-training (VLP) models struggle with category coupling and scale variation.
- These limitations hinder performance on recognizing unseen image categories.
Purpose of the Study:
- To propose a novel method, Hierarchical Dual-Granularity Alignment-CLIP (HDG-CLIP), for OV-MLIC.
- To enhance cross-category knowledge transfer by addressing category coupling and scale variation.
- To improve the recognition of unseen categories in images.
Main Methods:
- Developed HDG-CLIP, a novel OV-MLIC method leveraging VLP models.
- Constructed semantic category prototypes to mitigate category coupling.
- Implemented a sample-category dual-granularity matching mechanism to address scale variation.
- Utilized interaction between visual embeddings and category prototypes for feature decoupling.
Main Results:
- HDG-CLIP demonstrates state-of-the-art performance on OV-MLIC tasks.
- The method effectively handles category coupling and scale variation.
- Significant improvements observed on NUS-WIDE and Open-Images datasets.
Conclusions:
- HDG-CLIP offers a significant advancement in Open-Vocabulary Multi-Label Image Classification.
- The proposed method successfully addresses key challenges in recognizing unseen categories.
- HDG-CLIP provides a robust framework for future research in image classification.
