Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals

Summary

This study introduces a Deep Semantic Multimodal Hashing Network (DSMHN) for efficient image-text and video-text retrieval. The DSMHN significantly improves retrieval accuracy by jointly learning hash functions that preserve semantic labels and intermodality similarities.