Related Experiment Video
Updated: Dec 1, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
8.0K
ALSA: Adversarial Learning of Supervised Attentions for Visual Question Answering.
IEEE Transactions on Cybernetics
|November 11, 2020
Summary
This study introduces a novel Visual Question Answering (VQA) model using adversarial learning of supervised attentions (ALSAs). ALSAs effectively integrates human attention priors for improved question-relevant image region focus, enhancing VQA performance.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Visual Question Answering (VQA) relies on attention mechanisms to link questions with relevant image regions.
- Existing VQA methods struggle with attention by either using limited region types or neglecting human attention priors.
- This limitation hinders accurate inference for questions involving foreground objects and background context.
Purpose of the Study:
- To propose a novel VQA model, Adversarial Learning of Supervised Attentions (ALSAs), that addresses limitations in current attention mechanisms.
- To incorporate prior knowledge of human attention into the VQA attention distribution learning process.
- To improve the VQA model's ability to focus on question-relevant image regions for more accurate answer inference.
Main Methods:
- Developed two supervised attention modules: free-form based and detection-based, to leverage prior knowledge.
- Implemented an adversarial learning mechanism for interplay between the two attention modules.
- Utilized adversarial learning to mutually reinforce modules, enhancing multi-view feature correlation for answer inference.
Main Results:
- The proposed ALSAs model demonstrated favorable performance across three common VQA datasets.
- The integration of supervised attention modules and adversarial learning improved focus on question-relevant image areas.
- The method effectively learned correlations between questions and images from both free-form and detection-based views.
Conclusions:
- ALSAs offers a significant advancement in VQA by effectively utilizing human attention priors.
- The adversarial learning framework enhances the robustness and accuracy of attention mechanisms in VQA.
- The model's superior performance validates the approach for complex visual question answering tasks.
Related Concept Videos
Observational Learning
682
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
682
Associative Learning
908
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
908
