Related Experiment Video
Updated: May 10, 2026

11:34
High-resolution, High-speed, Three-dimensional Video Imaging with Digital Fringe Projection Techniques
Published on: December 3, 2013
TDP-DETR: Temporal dynamics perception framework for video moment retrieval and highlight detection
Huilin An1, Zefan Zhang2, Shijie Jiang2
1College of Software,Jilin University, Changchun, 130012, China.
Summary
This study introduces a new Transformer model for video analysis, improving temporal boundary detection in video moment retrieval and highlight detection. The method enhances action semantic alignment and reduces boundary bias in dynamic video content.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Video Moment Retrieval (VMR) and Highlight Detection (HD) are crucial for analyzing untrimmed videos.
- Accurate temporal boundary perception is essential for VMR and HD tasks.
- Current models struggle with precise action semantic alignment and exhibit boundary perception bias due to dynamic visual cues.
Purpose of the Study:
- To propose a novel Temporal Dynamics Perception DEtection TRansformer (TDP-DETR) for improved VMR and HD.
- To address limitations in current models regarding temporal dynamics and boundary perception.
- To enhance action semantic alignment and reduce boundary bias in video analysis.
Main Methods:
- TDP-DETR models action temporal dynamics from temporal persistence and temporal progression perspectives.
- A dynamic masking strategy is used for action duration-aware temporal modeling, incorporating temporal priors.
- An action state difference perception module captures frame-to-frame variations to perceive action progression speed.
Main Results:
- The proposed TDP-DETR method consistently outperforms existing state-of-the-art approaches on three MR/HD benchmarks.
- The approach demonstrates improved temporal boundary perception and action semantic alignment.
- The dynamic masking and action state difference modules contribute to enhanced performance.
Conclusions:
- TDP-DETR effectively models action temporal dynamics for precise boundary prediction in VMR and HD.
- The method offers a significant advancement in handling temporally dynamic video content.
- Future work will involve making the code publicly available for further research.

