Related Experiment Video
Updated: May 24, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
8.9K
Exploiting Unlabeled Videos for Video-Text Retrieval via Pseudo-Supervised Learning.
Summary
This study introduces Pseudo-Supervised Selective Contrastive Learning (PS-SCL) for video-text retrieval, reducing reliance on manual annotations. PS-SCL effectively trains models using automatically generated pseudo-texts and selective contrastive learning, improving performance on benchmarks.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Large-scale pre-trained vision-language models like CLIP excel at video-text retrieval (VTR).
- Traditional VTR methods require costly, labor-intensive manual annotation of video-text pairs.
- Existing techniques often fine-tune models directly on clean, annotated data, limiting scalability.
Purpose of the Study:
- To develop a novel approach for video-text retrieval that minimizes dependency on manual text annotations.
- To leverage unlabeled video data for training more efficient and scalable VTR models.
- To enhance multi-modal learning under weak supervision conditions.
Main Methods:
- Introduced Pseudo-Supervised Selective Contrastive Learning (PS-SCL) to generate pseudo-supervisions from unlabeled video data.
- Utilized CLIP's visual recognition to automatically generate pseudo-texts, providing weak textual guidance.
- Developed Selective Contrastive Learning (SeLeCT) to prioritize and select highly correlated pseudo-supervised video-text pairs for effective multi-modal learning.
Main Results:
- PS-SCL significantly outperforms CLIP's zero-shot performance across multiple video-text retrieval benchmarks.
- Achieved notable improvements, including 8.2% R@1 on MSRVTT, 12.2% R@1 on DiDeMo, and 10.9% R@1 on ActivityNet for video-to-text retrieval.
- Demonstrated the effectiveness of pseudo-supervision and selective contrastive learning in weak pairing scenarios.
Conclusions:
- PS-SCL offers a scalable and effective alternative to traditional VTR methods that rely on extensive manual annotations.
- The proposed approach successfully bridges the gap between visual content and textual descriptions using weak supervision.
- This work advances the field of video-text retrieval by enabling robust multi-modal learning with reduced annotation costs.

