Related Experiment Video
Updated: Oct 30, 2025

Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Linguistic issues behind visual question answering
Raffaella Bernardi1, Sandro Pezzelle2
1CIMeC and DISI University of Trento Trento Italy.
Answering questions grounded in images requires understanding language and visuals. While progress is made, a unified computational approach is still needed for visual question answering.
Area of Science:
- Computational Linguistics
- Computer Vision
- Cognitive Science
Background:
- Visually-grounded question answering (VQA) is a complex task involving linguistic and visual understanding.
- Historically, VQA has been a challenge for computational natural language understanding (NLU) systems.
- Recent machine learning advancements have revitalized VQA research at the intersection of AI fields.
Purpose of the Study:
- To review current approaches to VQA, including datasets, models, and frameworks.
- To analyze VQA progress from a theoretical linguistics perspective using established desiderata.
- To identify gaps and propose future research directions for a unified VQA approach.
Main Methods:
- Literature review of VQA research.
- Analysis of computational achievements against theoretical linguistic desiderata.
- Synthesis of current trends and identification of future research needs.
Main Results:
- Significant progress has been achieved in VQA, reconciling engineering and theoretical perspectives.
- Current VQA systems demonstrate impressive capabilities in integrating visual and linguistic information.
- However, a comprehensive approach addressing all linguistic challenges in VQA remains an open problem.
Conclusions:
- Further research is essential to develop a unified computational framework for VQA.
- Future work should focus on integrating semantic, syntactic, and pragmatic understanding with visual context.
- The field needs to bridge the gap between theoretical linguistic requirements and current AI capabilities for VQA.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Related Concept Videos
Language and Cognition
Components of Language
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Tip-of-the-Tongue Phenomenon
Detection of Gross Error: The Q Test
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...