Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Probabilistic Embeddings With Evidence Learning and Refinement for Text-Video Retrieval
Summary
This study introduces Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR) for improved text-video retrieval. PE2LR enhances cross-modal alignment by modeling video-text pairs as probability distributions, achieving state-of-the-art performance.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Text-video retrieval faces challenges due to semantic gaps between modalities.
- Ambiguity arises from differences in information granularity, abstraction levels, and temporal variations.
- Existing methods struggle with reliable sample-level alignment.
Purpose of the Study:
- To develop a novel method for accurate cross-modal alignment in text-video retrieval.
- To address the inherent ambiguity and uncertainty in matching heterogeneous video and text data.
- To improve the performance of text-video search systems.
Main Methods:
- Proposed Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR) method.
- Modeled video-text pairs as probability distributions using evidence theory.
- Implemented distribution-level representation learning and a distribution-based embedding refinement module.
Main Results:
- PE2LR effectively resolves semantic ambiguity between video and text pairs.
- The method enhances semantic consistency across modalities.
- Achieved state-of-the-art search performance on benchmark datasets (MSRVTT, DiDeMo, ActivityNet Captions).
Conclusions:
- PE2LR offers a robust approach to text-video retrieval by handling uncertainty.
- The probabilistic modeling and refinement strategy significantly improve cross-modal alignment.
- The method demonstrates superior performance over existing techniques.