Related Experiment Videos
Statistical Model-Driven Similarity Hashing for Unsupervised Multimedia Retrieval
Abstract:
Unsupervised deep cross-modal hash retrieval aims to map multi-modal features into binary hash codes without labels, which is of interest due to the storage efficiency, query speed, and convenient applications. However, existing approaches suffer from two main limitations: (1) Insufficient consideration of text instance similarity, along with independent or redundant fusion to learn multi-modal similarity information. (2) Ignoring the noisy adjacent correlations between multi-modal instances leads to a lack of discriminative capacity in the generated hash codes. To address the challenges, we propose a novel approach called Statistical Model-driven Similarity Hashing. Specifically, Jaccard similarity is introduced to construct the text similarity matrix, which reduces the similarity error between text instances while better considering the asymmetry of elements in text features. After that, original similarity information between various modalities is integrated to construct a unified similarity matrix, where the modality gaps are effectively bridged while reducing the redundant information. In addition, we present a Statistical Model-driven Similarity Enhancement approach, which reduces the noise of similarity relations between multi-modal instances by utilizing Gaussian Mixture Models to keep instances with lower semantic similarity as far away from each other as possible. Moreover, we introduce an enhanced version, which integrates a structure-aware self-paced contrastive learning framework that progressively schedules reliable and fuzzy-positive pairs, jointly optimizing statistical reconstruction and discriminative alignment, and further extends to noisy environments. Extensive experiments on three benchmark datasets demonstrate the excellent performance of the proposed method.