Related Experiment Video
Updated: Jan 14, 2026

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
Causality guided co-attention network for visual question answering
Jiali Miao1, Kui Yu1, Baofu Fang2
1Key Laboratory of Knowledge Engineering with Big Data (the Ministry of Education of China), Hefei University of Technology, Hefei, 230601, China; School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, 230601, China.
Abstract:
Visual Question Answering (VQA) answers image-related textual questions by analyzing and learning multimodal representations of visual images and language. With the popularity of multimodal large-scale language modeling research, VQA has become a paradigm task for large models. However, when numerous overlapping visual objects are present, the self-attention mechanism struggles to focus on relevant ones. This results in suboptimal visual representations that hinder accurate question answering in VQA models. Conversely, relevant object relations tend to be sparse, exclusively learning relevant objects to represent features may be overly restrictive and detrimental to performance. In practical applications, appropriate background and spurious features can improve model performance, like "Chopsticks" in prediction for "Noodle Soup". Inspired by this, we propose the Causality Guided Co-attention Network (CGCN), a deep hierarchical granular visual feature attention network. Specifically, we first introduce a causal graph to model relations of region object features and divide the object features into two groups: the Influential Object Features (IOFs) and Supportive Object Features (SOFs), which the IOFs denote the object features that are strongly relevant or crucial to the target object and the SOFs represent the objects that are weakly relevant to the target object in background or spurious features. Then, we propose a new causality inspired co-attention network, which utilizes causal graphs to guide the self-attention mechanism learning cross-modal representations of the IOFs and SOFs with question features. Finally, we design a dual output multi-modal feature fusion method to promote the model performance. The experimental results show that the VQA performance of CGCN is significantly improved. Our codes are available at https://github.com/Miao-jiaLi/CGCN.
Related Concept Videos
The Anchoring-and-Adjustment Heuristic
Associative Learning
Classical conditioning, also known...
Observational Learning
Visual Agnosia
Visual System
Once through the pupil, the light passes through the lens, a...