Related Experiment Video
Updated: Aug 8, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Lightweight dense video captioning with cross-modal attention and knowledge-enhanced unbiased scene graph
Shixing Han1, Jin Liu1, Jinyingming Zhang1
1College of Information Engineering, Shanghai Maritime University, Shanghai, 201306 China.
This study introduces CMCR, a novel dense video captioning model that integrates audio-visual information and commonsense reasoning for more accurate scene descriptions and event localization. Experiments show CMCR outperforms existing methods.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Dense video captioning (DVC) traditionally focuses on visual features, often neglecting audio cues.
- This oversight leads to inaccuracies in locating events within video scenes.
- Existing DVC models struggle with overlapping events and logical caption coherence.
Purpose of the Study:
- To develop a novel DVC model (CMCR) that effectively integrates multi-modal information (audio-visual).
- To improve the accuracy of scene event localization and caption generation.
- To enhance the logical consistency and reasoning capabilities of video descriptions.
Main Methods:
- Proposed a Cross-Modal processing (CM) module using cross-modal attention for multi-modal feature encoding.
- Introduced an event refactoring algorithm to address inaccurate localization of overlapping events.
- Incorporated a Commonsense Reasoning (CR) module with a knowledge-enhanced scene graph for improved caption logic.
Main Results:
- The CMCR model demonstrated superior performance compared to state-of-the-art methods on the ActivityNet Captions dataset.
- Ablation studies confirmed the significant contributions of both the CM and CR modules.
- The model achieved more accurate event localization and logically coherent captions.
Conclusions:
- Integrating audio-visual information and commonsense reasoning significantly enhances dense video captioning.
- The proposed CMCR model offers a robust solution for accurate and contextually relevant video scene descriptions.
- Future work can explore further refinements in multi-modal fusion and knowledge integration for DVC.
More Related Videos
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
Related Concept Videos
Stereotype Content Model
Observational Learning
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
The Anchoring-and-Adjustment Heuristic
Nonconscious Mimicry
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...