Related Experiment Video
Updated: Apr 20, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
V-Sparse: From temporal-spatial visual semantic compression to coarse-to-fine interaction for text-video retrieval.
Xin Liu1, Shibai Yin1, Jun Wang2
1School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics, Chengdu, Sichuan,610000, China; Engineering Research Center of Intelligent Finance, Ministry of Education, Chengdu, Sichuan, 610000, China.
The V-Sparse model enhances text-to-video retrieval by using visual semantic compression and coarse-to-fine alignment. This approach improves cross-modal matching accuracy for better video search results.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Text-to-video retrieval aims to find relevant videos using text queries from large unlabeled collections.
- CLIP-based methods focus on feature enhancement and interaction but struggle with cross-modal matching imbalance.
- Current methods are insufficient for high-precision retrieval due to modality differences.
Purpose of the Study:
- To propose a novel text-video retrieval model, V-Sparse, addressing cross-modal matching imbalance.
- To introduce visual semantic compression (VSC) for feature enhancement and coarse-to-fine alignment (CFI) for feature interaction.
- To improve the precision and effectiveness of text-to-video retrieval.
Main Methods:
- Developed a text-guided Visual Semantic Compression (VSC) module with Temporal (TVSC) and Spatial (SVSC) components to reduce feature redundancy.
- Introduced a Coarse-to-Fine granularity Interaction (CFI) module for aligning sentences with frames, sentences with patches, and words with patches.
- Employed a unified joint feature encoding perspective for VSC and CFI to facilitate cross-modal alignment.
Main Results:
- V-Sparse achieved state-of-the-art results on six benchmark datasets for both long-video and short-text retrieval.
- Demonstrated the effectiveness of feature compression in cross-modal interaction through extensive ablation studies.
- Showcased V-Sparse as an effective intermediate pathway for modality interaction.
Conclusions:
- V-Sparse effectively mitigates the inherent imbalance in text-video modal pairing.
- The proposed VSC and CFI modules significantly enhance cross-modal text-video alignment.
- V-Sparse offers a promising direction for advancing text-to-video retrieval technologies.
Related Concept Videos
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Chunking and Rehearsal in Sensory Memory
Depth Perception and Spatial Vision
Parallel Processing
Visual System
Once through the pupil, the light passes through the lens, a...
Storage