Related Experiment Video
Updated: Jul 24, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Fusion of Multi-Modal Features to Enhance Dense Video Caption
Xuefei Huang1, Ka-Hou Chan1,2, Weifan Wu1
1Faculty of Applied Sciences, Macao Polytechnic University, Macau 999078, China.
This study introduces a novel fusion model for dense video captioning, integrating both visual and audio features. The model enhances video analysis by combining Transformer and LSTM frameworks, achieving competitive results.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Dense video captioning aims to generate descriptive text for video content.
- Existing methods often overlook crucial audio information, relying solely on visual features.
- This limitation hinders comprehensive video content analysis.
Purpose of the Study:
- To develop an advanced fusion model for dense video captioning.
- To integrate both visual and audio features for more accurate video understanding.
- To improve the performance of video captioning systems.
Main Methods:
- A fusion model employing the Transformer framework to integrate visual and audio features.
- Utilizing multi-head attention to manage varying sequence lengths.
- Implementing a Common Pool for feature alignment, redundancy filtering, and confidence scoring.
- Employing an LSTM (Long Short-Term Memory) network as a decoder for sentence generation.
Main Results:
- The proposed fusion model demonstrates competitive performance on the ActivityNet Captions dataset.
- Integration of audio features alongside visual features enhances captioning accuracy.
- The model effectively handles variations in sequence lengths and filters redundant information.
Conclusions:
- The fusion model offers a significant advancement in dense video captioning.
- Combining visual and audio data provides a more holistic approach to video analysis.
- The method shows promise for real-world applications requiring detailed video understanding.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence...
Tagging and Fusion Proteins
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Extraction: Advanced Methods
Upsampling

