Related Experiment Video
Updated: Jun 29, 2025

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
3.7K
Latent Attention Network With Position Perception for Visual Question Answering
IEEE Transactions on Neural Networks and Learning Systems
|March 26, 2024
Summary
This study introduces a novel latent attention network to improve visual question answering (VQA) by better understanding object positions. The new method enhances performance on complex VQA tasks involving multiple objects and spatial relationships.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Visual Question Answering (VQA) models struggle with complex spatial relationships between multiple objects.
- Understanding relative positions described by prepositions is crucial for accurate VQA.
Purpose of the Study:
- To develop a novel latent attention (LA) network for VQA that effectively addresses complex relative position relationships.
- To improve the model's ability to comprehend multiobject correlations and enhance performance on counting questions.
Main Methods:
- Proposed a latent attention generation module (LAGM) to reconstruct attention based on positional prepositions.
- Introduced a position-aware module (PAM) to encode absolute and relative position relations.
- Developed a gated counting module (GCM) to improve quantitative reasoning for counting questions.
Main Results:
- The proposed LA network accurately captures complex relative position features, guiding attention to correct objects/regions.
- The PAM enhances comprehension of multiobject correlations.
- The GCM improves performance on counting questions.
Conclusions:
- The novel LA network significantly improves VQA performance, particularly in scenarios with complex spatial reasoning.
- The method outperforms state-of-the-art approaches on widely used VQA datasets (VQA v1 and v2).

