Related Experiment Video
Updated: Sep 14, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
MGTR-MISS: More Ground Truth Retrieving based Multimodal Interaction and Semantic Supervision for video description
Jiayu Zhang1, Pengjie Tang1, Yunlan Tan1
1Key Laboratory of Electronic Data Control and Evidence Collection in Jiangxi Province, Jinggangshan University, Ji'an 343009, PR China; College of Electronics & Information Engineering, Jinggangshan University, Ji'an 343009, PR China.
This study introduces MGTR-MISS, a novel model for generating accurate video descriptions by integrating visual and linguistic information. It significantly improves video captioning performance on benchmark datasets.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Generating accurate video descriptions is challenging due to the need for effective multimodal interaction.
- Current models often fail to adequately align visual and linguistic modalities.
Purpose of the Study:
- To propose a novel model, MGTR-MISS, for generating more accurate and semantically rich video descriptions.
- To enhance video captioning by effectively integrating external language knowledge and improving multimodal alignment.
Main Methods:
- MGTR-MISS utilizes multimodal interaction and semantic supervision.
- External language knowledge is retrieved from training data to enrich linguistic semantics.
- A multimodal interaction module aligns visual and linguistic features, followed by a caption generator with visual-textual attention.
Main Results:
- MGTR-MISS outperforms baseline and state-of-the-art methods on MSVD, MSR-VTT, and VATEX datasets.
- Achieved CIDEr scores of 111.1 on MSVD and 55.0 on MSR-VTT.
Conclusions:
- The proposed MGTR-MISS model demonstrates superior performance in video description generation.
- Effective multimodal interaction and semantic supervision are key to improving video captioning accuracy and richness.

