Related Experiment Video
Updated: Jul 9, 2025

07:36
Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
15.7K
A cross-modal conditional mechanism based on attention for text-video retrieval
Wanru Du1,2, Xiaochuan Jing2, Quan Zhu1,2
1China Aerospace Academy of Systems Science and Engineering, Beijing 100048, China.
Mathematical Biosciences and Engineering : MBE
|December 5, 2023
Summary
This study introduces a novel cross-modal model for text-video retrieval, focusing on detailed frame-text alignment. The proposed method enhances matching accuracy by considering both global topics and specific details for improved video understanding.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Current cross-modal retrieval often aligns only global video and sentence features.
- Videos contain richer information than text, necessitating finer-grained matching.
- Existing methods may miss critical frame-level details crucial for accurate text-video correspondence.
Purpose of the Study:
- To develop an advanced cross-modal model for text-video retrieval.
- To improve matching by focusing on frame-level semantics and detailed information.
- To enhance the understanding of video content through text-driven feature extraction.
Main Methods:
- Proposed a cross-modal conditional feature aggregation model utilizing an attention mechanism.
- Introduced a cross-modal attentional feature aggregation module for extracting relevant frame features guided by text semantics.
- Implemented a global-local similarity calculation module for dual-granularity matching (video-sentence and frame-word).
Main Results:
- The model demonstrated superior performance over state-of-the-art methods on four benchmark datasets (MSR-VTT, LSMDC, MSVD, DiDeMo).
- The cross-modal attention aggregation effectively captured primary semantic video information.
- The global-local similarity calculation accurately matched text and video based on both topic and detail features.
Conclusions:
- The proposed model significantly advances text-video retrieval capabilities.
- Attention-based feature aggregation is effective for capturing salient video semantics.
- Dual-granularity similarity calculation enhances the precision of text-video matching.

