Related Experiment Video
Updated: Aug 5, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Enhancing Image-Text Retrieval via Region-Grid Interaction and Semantic Calibration
1School of Mechanical and Power Engineering, Nanjing Tech University, Nanjing 211816, China.
Sensors (Basel, Switzerland)
|July 28, 2026
Summary
This study introduces a novel network for image-text retrieval, enhancing semantic alignment by integrating region and grid features. The proposed method improves retrieval accuracy by effectively combining object-centric and contextual visual information.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Accurate semantic alignment between visual content and text is crucial for image-text retrieval.
- Existing methods often rely on object-centric region features, potentially missing contextual details.
- Grid features offer dense spatial information but lack explicit semantic structure.
Purpose of the Study:
- To propose a novel Region-Grid Interaction and Calibration Network (RGICN) for improved image-text retrieval.
- To effectively combine complementary region and grid features for better visual representation.
- To enhance the discriminative power of visual embeddings for retrieval tasks.
Main Methods:
- A Global-Guided Feature Interaction Module exchanges information between region and grid features using global semantics.
- A Text-Guided Feature Calibration Module refines visual features by suppressing irrelevant content using auxiliary descriptions.
- An Adaptive Gating Fusion Module dynamically integrates diverse visual representations.
Main Results:
- The RGICN effectively integrates object-level semantics and contextual cues.
- The network calibrates visual features, suppressing text-irrelevant information.
- Experiments on MS-COCO and Flickr30K datasets show competitive performance against state-of-the-art methods.
Conclusions:
- The proposed RGICN model significantly improves image-text retrieval performance.
- Integrating region and grid features through interaction and calibration enhances semantic alignment.
- The RGICN offers a more comprehensive and discriminative visual embedding for retrieval.