Related Experiment Video
Updated: Dec 22, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
8.0K
SMAN: Stacked Multimodal Attention Network for Cross-Modal Image-Text Retrieval
IEEE Transactions on Cybernetics
|May 10, 2020
Summary
This study introduces a stacked multimodal attention network (SMAN) for improved cross-modal image-text retrieval. The SMAN precisely models fine-grained image-text correlations, enhancing retrieval accuracy.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Cross-modal image-text retrieval is challenging due to limitations in existing global and local representation alignment methods.
- Global methods struggle to identify semantically relevant image and text portions.
- Local methods face computational burdens when aggregating fragment similarities.
Purpose of the Study:
- To propose a novel Stacked Multimodal Attention Network (SMAN) for fine-grained cross-modal retrieval.
- To improve the precision of image-text similarity measurement by exploiting interdependencies.
- To address the computational complexity and semantic pinpointing issues in current retrieval techniques.
Main Methods:
- Developed a Stacked Multimodal Attention Network (SMAN) utilizing a stacked multimodal attention mechanism.
- Employed sequential intramodal and multimodal information for multi-step attention reasoning.
- Introduced a novel bidirectional ranking loss to preserve data manifold structure.
Main Results:
- The SMAN effectively models fine-grained correlations between image regions and text words.
- The network precisely identifies semantically meaningful visual and textual elements for similarity measurement.
- Achieved competitive performance on benchmark datasets compared to state-of-the-art methods.
Conclusions:
- The proposed SMAN offers a more precise and computationally efficient approach to cross-modal image-text retrieval.
- The attention mechanism and bidirectional loss enhance the modeling of intermodal relationships.
- SMAN demonstrates superior performance, advancing the field of multimodal understanding.
