Related Experiment Videos
TIC: Co-optimization of training and inference for efficient image-text retrieval
Zhongxiang You1, Shiyuan Yang1, Zhixin Li1
1Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University, Guilin, 541004, China; Guangxi Key Lab of Multi-source Information Mining and Security, Guangxi Normal University, Guilin, 541004, China.
Abstract:
Embedding-based image-text retrieval methods achieve efficient inference by independently encoding different modalities into a shared space. However, their performance is constrained by a dual limitation: fragile representations during training due to insufficient exploration of the semantic manifold, and biased similarity metrics during inference caused by the "hubness" problem. To address these challenges, we propose a Training-Inference Co-optimization (TIC) framework that systematically enhances both stages of the retrieval pipeline. During training, we introduce Embedding Space Self-Expansion (ES2)-a novel augmentation strategy that generates multiple dynamic views of image representations via random attention masking. This effectively transforms static sample points into local probabilistic spaces, enabling the model to explore broader embedding boundaries and learn more robust features. We further incorporate pre-trained GloVe semantic priors to initialize the text encoder, thereby providing linguistic knowledge that improves cross-modal alignment. During inference, we apply Dual Softmax Calibration-a zero-parameter post-processing technique that recalibrates the similarity matrix by incorporating global distribution information to mitigate hubness artifacts. Extensive experiments on the MS-COCO and Flickr30K benchmarks demonstrate the effectiveness of our approach. Our method achieves state-of-the-art performance among embedding-based techniques, significantly outperforming strong baseline models-even when utilizing a more lightweight text encoder.