Decoupled Cross-Modal Phrase-Attention Network for Image-Sentence Matching

Summary

This study introduces a new Decoupled Cross-modal Phrase-Attention network (DCPA) for better image-sentence matching. It models phrase relationships, improving retrieval accuracy over word-level alignments.