Related Experiment Video
Updated: Jul 26, 2025

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
Efficient Token-Guided Image-Text Retrieval With Consistent Multimodal Contrastive Training
This study introduces a unified framework for image-text retrieval, combining coarse- and fine-grained representations. The Token-Guided Dual Transformer (TGDT) architecture improves retrieval accuracy and efficiency.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Image-text retrieval is crucial for understanding semantic relationships between visual and linguistic data.
- Existing methods often focus on either global or local features, neglecting their interplay, leading to suboptimal accuracy and high computational costs.
Purpose of the Study:
- To develop a novel framework that integrates coarse- and fine-grained representation learning for enhanced image-text retrieval.
- To improve retrieval accuracy and reduce computational complexity in multimodal understanding tasks.
Main Methods:
- Proposed the Token-Guided Dual Transformer (TGDT) architecture with two homogeneous branches for image and text processing.
- Introduced a Consistent Multimodal Contrastive (CMC) loss to ensure semantic consistency across modalities in a shared embedding space.
- Implemented a two-stage inference method utilizing mixed global and local cross-modal similarity.
Main Results:
- Achieved state-of-the-art retrieval performance on benchmark datasets.
- Demonstrated significantly lower inference time compared to existing representative methods.
- The unified framework effectively leverages both coarse- and fine-grained information.
Conclusions:
- The proposed TGDT architecture offers a more effective and efficient approach to image-text retrieval by unifying multimodal representations.
- The CMC loss and two-stage inference method contribute to superior semantic understanding and retrieval accuracy.
- This work provides a new perspective on multimodal learning, aligning with human cognitive processes.
More Related Videos
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025