Related Experiment Video
Updated: Dec 19, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
917
Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals
Summary
This study introduces a Deep Semantic Multimodal Hashing Network (DSMHN) for efficient image-text and video-text retrieval. The DSMHN significantly improves retrieval accuracy by jointly learning hash functions that preserve semantic labels and intermodality similarities.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Hashing is crucial for efficient large-scale multimedia data retrieval.
- Existing methods face challenges in scalable multimodal retrieval.
Purpose of the Study:
- To propose a novel Deep Semantic Multimodal Hashing Network (DSMHN) for scalable image-text and video-text retrieval.
- To develop a unified framework that learns compact and high-quality hash codes.
Main Methods:
- Leveraging 2-D and 3-D Convolutional Neural Networks (CNNs) for spatial and spatio-temporal feature extraction.
- Jointly learning modality-specific hash functions by preserving intermodality similarities and intramodality semantic labels.
- Employing a unified framework integrating feature representation, similarity-preserving, semantic label-preserving, and hash function learning.
Main Results:
- The DSMHN framework demonstrates superior performance in both single-modal and cross-modal retrieval tasks.
- Experimental results on four benchmark datasets show significant improvements over state-of-the-art methods.
- The proposed method achieves high accuracy in image-text and video-text retrieval.
Conclusions:
- The DSMHN is a generic, scalable, and flexible deep hashing framework for multimodal retrieval.
- The method effectively learns compact and discriminative hash codes for efficient searching.
- DSMHN offers a significant advancement in scalable multimodal retrieval systems.
