A cross-modal conditional mechanism based on attention for text-video retrieval

Wanru Du1,2, Xiaochuan Jing2, Quan Zhu1,2

  • 1China Aerospace Academy of Systems Science and Engineering, Beijing 100048, China.

Summary

This study introduces a novel cross-modal model for text-video retrieval, focusing on detailed frame-text alignment. The proposed method enhances matching accuracy by considering both global topics and specific details for improved video understanding.

Related Concept Videos