Related Experiment Video
Updated: Jan 9, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
RCVQA: Visual question answering model based on reading comprehension
Deguang Chen1, Jianrui Chen2, Zhongshi Shao1
1School of Artificial Inteligence and Computer Science, Shaanxi Normal University, Xi'an, 710119, Shaanxi, China.
None:
Visual Question Answering (VQA) combines computer vision and natural language processing to answer image-related questions. Current models have three shortcomings: (1) Their limited interaction with other established fields hinders deep semantic analysis of both images and questions, restricting practical applications; (2) Predicted answers are often limited to single words or short phrases, lacking diversity and completeness to meet varied user needs; (3) Exclusive reliance on accuracy oversimplifies evaluation, compromising comprehensive performance assessment and model optimization. To address these challenges, we propose RCVQA, a novel reading-comprehension-based VQA model that introduces key innovations in methodology and evaluation metrics. First, dataset text paragraphs are preprocessed to remove irrelevant information, enhancing contextual focus. To ensure comprehensive evaluation, three model variants-RCVQAP, RCVQAC, and VQAT-are designed and tested across four popular datasets. To address the limitations of single-answer prediction in existing methods, two innovative algorithms are developed: answer position prediction and answer content prediction, enabling multi-span answer retrieval for more comprehensive reasoning. Recognizing the limitations of accuracy as a sole evaluation metric, four novel metrics-PPR∥, TPR∥, H∥-Means, and ESM∥-are introduced to provide more comprehensive criteria for evaluating the RCVQA system. Additionally, we integrate strategies such as image captioning and local training, enhancing content understanding and ensuring data security. Experimental results demonstrate that RCVQA significantly outperforms state-of-the-art methods, achieving 1 %-7 % improvements across benchmarks: A-OKVQA (DA:+1.05 %, MC:+2.22 %), KR-VQA (+7.28 %), GQA (Val:+7.10 %, Test:+7.06 %), and OK-VQA (+4.30 %). Our implementation codes will be publicly available at: https://github.com/jianruichen/RCVQA.
More Related Videos
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
05:54Eye-tracking to Distinguish Comprehension-based and Oculomotor-based Regressive Eye Movements During Reading
Published on: October 18, 2018
Related Concept Videos
Visual System
Once through the pupil, the light passes through the lens, a...
Vision
Visual Agnosia
Information Processing Approach
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...