Related Experiment Video
Updated: Aug 1, 2025

A Two-interval Forced-choice Task for Multisensory Comparisons
Published on: November 9, 2018
NExT-OOD: Overcoming Dual Multiple-Choice VQA Biases.
This study introduces NExT-OOD, a benchmark for video question answering (videoQA), to address biases in multimodal AI models. A novel graph-based method effectively reduces these biases, improving model generalization and reasoning abilities.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Multiple-choice Visual Question Answering (VQA) models show progress but rely on dataset correlations, limiting multimodal understanding and generalization.
- Spurious correlations, specifically Vision-Answer (VA) bias and Question-Answer (QA) bias, hinder VQA model performance.
Purpose of the Study:
- To systematically investigate VA and QA biases in VQA models.
- To develop a benchmark for evaluating VQA model generalizability and reasoning in out-of-distribution settings.
- To propose a novel method for reducing identified biases in VQA models.
Main Methods:
- Construction of the NExT-OOD benchmark with three sub-datasets (NExT-OOD-VA, NExT-OOD-QA, NExT-OOD-VQA) for bias quantification.
- Evaluation of existing VQA models on NExT-OOD to demonstrate performance degradation.
- Proposal of a graph-based cross-sample method with a contrastive graph matching loss for bias reduction.
Main Results:
- Existing VQA models exhibit significant performance degradation on the NExT-OOD benchmark compared to standard datasets.
- The proposed graph-based method effectively mitigates VA and QA biases by leveraging cross-sample information.
- The approach encourages models to focus on multimodal content rather than spurious correlations.
Conclusions:
- The NExT-OOD benchmark provides a scientific tool for assessing VQA model generalizability and reasoning.
- The proposed bias reduction method significantly outperforms existing strategies.
- The developed approach demonstrates effectiveness and generalizability for improving VQA models.
More Related Videos
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
Related Concept Videos
Confirmation Biases
Hindsight Biases
The Anchoring-and-Adjustment Heuristic
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
The Representativeness Heuristic
Visual Agnosia