Related Experiment Video
Updated: Dec 10, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
876
An Effective Dense Co-Attention Networks for Visual Question Answering
1College of Information Engineering, Shanghai Maritime University, Shanghai 201306, China.
Sensors (Basel, Switzerland)
|September 3, 2020
Summary
Dense Co-Attention Networks (DCAN) improve visual question answering by enhancing multimodal interactions. This dense co-attention model boosts accuracy for AI applications.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Current Visual Question Answering (VQA) methods primarily use co-attention for coarse multimodal interactions.
- Existing approaches often overlook dense self-attention mechanisms within the question modality.
- This limitation hinders the fine-grained understanding required for complex VQA tasks.
Purpose of the Study:
- To propose an effective Dense Co-Attention Networks (DCAN) model for improved VQA accuracy.
- To address the limitations of current VQA models by incorporating dense self-attention.
- To enhance the fine-grained interaction between textual queries and visual elements.
Main Methods:
- Utilized Bidirectional Long Short-Term Memory (Bi-LSTM) for robust encoding of questions and answers, capturing long-range dependencies.
- Introduced a dense multimodal co-attention model for fine-grained interactions between question words and image regions.
- Designed a hierarchical structure by cascading self-attention and guided-attention units.
Main Results:
- The proposed DCAN model demonstrated significant performance advantages on the VQA-v2 dataset.
- DCAN achieved higher accuracy compared to state-of-the-art VQA approaches.
- The model effectively captures complex relationships between visual objects and textual queries.
Conclusions:
- Dense Co-Attention Networks (DCAN) offer a superior approach to Visual Question Answering.
- The model's enhanced interaction mechanisms lead to improved accuracy and robustness.
- DCAN broadens the applicability of VQA in diverse artificial intelligence scenarios.

