Related Experiment Video
Updated: Sep 1, 2025

07:36
Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
15.8K
Decoupled Cross-Modal Phrase-Attention Network for Image-Sentence Matching
Summary
This study introduces a new Decoupled Cross-modal Phrase-Attention network (DCPA) for better image-sentence matching. It models phrase relationships, improving retrieval accuracy over word-level alignments.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Current image-sentence matching methods often focus on word-level alignments.
- These approaches overlook the importance of phrase-level correspondences between visual and textual data.
- This limitation hinders the accurate understanding of complex relationships in multimodal data.
Purpose of the Study:
- To propose a novel network, the Decoupled Cross-modal Phrase-Attention (DCPA) network.
- To model relationships between textual and visual phrases for improved image-sentence matching.
- To develop a decoupled training and inference strategy to enhance bi-directional retrieval.
Main Methods:
- Developed a Decoupled Cross-modal Phrase-Attention (DCPA) network.
- Modeled interactions between visual phrases and textual phrases.
- Implemented a decoupled training and inference approach for separate optimization of image-to-sentence and sentence-to-image retrieval.
Main Results:
- Achieved state-of-the-art performance on Flickr30K and MS-COCO datasets.
- Demonstrated significant improvements in image-sentence matching accuracy compared to existing methods.
- Showcased competitive results, even when compared to methods utilizing external knowledge.
Conclusions:
- Phrase-level modeling is crucial for effective image-sentence matching.
- The proposed DCPA network offers a superior approach to multimodal understanding.
- The decoupled training strategy effectively addresses the bi-directional retrieval trade-off.

