Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Dynamic scale position embedding for cross-modal representation learning
Jungkyoo Shin1, Sungmin Kang1, Yoonsik Cho1
1Department of Artificial Intelligence, Chung-Ang University, 84, Heukseok-ro, Dongjak-gu, Seoul, Republic of Korea, Seoul, 06974, Seoul, Korea.
None:
In this paper, we introduce a novel approach to capture temporal information in videos across multiple scales for cross-modal learning. As videos naturally encapsulate semantic information of diverse durations, existing methods that primarily depend on fine- and coarse-grained contrastive learning may fail to fully capture the inherent semantic information. To bridge this gap, we propose Dynamic Scale Position Embedding (DSPE), a novel approach that enables a single transformer to interpret videos at various temporal scales through dynamic adjustment of temporal position embedding. In contrast to conventional multi-scale methods that aggregate video clips, DSPE maintains the distinct features of each clip, thus preserving semantic integrity and enhancing semantic content comprehension. Based on this, we present an efficient multi-scale temporal encoder designed to adeptly capture temporal information across a broad spectrum from fine to coarse granularity. Comprehensive experiments across four datasets-MSR-VTT, LSMDC, MSVD, and ActivityNet-Captions-and two distinct tasks-text-video retrieval and video-captioning-with consistent performance improvements highlight the significance of the presented multi-scale approach.
Related Concept Videos
Position and Displacement Vectors
Further, several important kinds of...
Position Vectors
For instance, we want to locate a point P(x, y, z) relative to the origin of coordinates O. In that case, we can define a position...
Scaling
Position and Displacement
Cross Product
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
Vector Transformation in Rotating Coordinate Systems

