Related Experiment Video
Updated: Sep 19, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Visual-Language Scene-Relation-Aware Zero-Shot Captioner
This study introduces a novel scene-relation-level pre-training task for zero-shot image captioning. The proposed Visual-Language Scene Relation Aware Captioner (SRACap) enhances image understanding and reduces caption hallucinations by focusing on scene relations.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Natural Language Processing
Background:
- Zero-shot image captioning leverages pre-trained visual language models (VLMs) and language models (LMs) for caption generation without paired training data.
- Existing methods focus on sentence-level or entity-level connections, but often suffer from hallucinations due to biased associations.
- There is a need for improved methods that can generate accurate and contextually relevant captions across different domains.
Purpose of the Study:
- To propose a novel scene-relation-level pre-training task for zero-shot image captioning.
- To introduce the Visual-Language Scene Relation Aware Captioner (SRACap) for improved image understanding and caption generation.
- To enhance cross-domain zero-shot generalization capabilities in image captioning.
Main Methods:
- Developed a scene-relation-level pre-training task, treating scene relations as key bridges between visual and textual modalities.
- Constructed SRACap, a model that predicts scene relations and generates captions, featuring a scene reinforcement switching pipeline for generalization.
- Employed a scene policy network for dynamic cropping of salient image regions and a mixture-of-rewards (MoR) module with expert CLIP models, optimized via policy gradient algorithm.
Main Results:
- SRACap demonstrates strong cross-domain zero-shot generalization capabilities.
- The model accurately understands scene structures and generates high-quality captions.
- Extensive experiments show SRACap significantly outperforms existing zero-shot inference methods on standard benchmarks.
Conclusions:
- The scene-relation-level pre-training approach effectively addresses limitations of previous methods in zero-shot image captioning.
- SRACap offers a robust solution for generating accurate, semantically consistent, and contextually relevant image captions.
- The proposed method advances the state-of-the-art in zero-shot image captioning, particularly in cross-domain generalization.
More Related Videos
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Related Concept Videos
Non-Verbal Cues
Visual Agnosia
Stereotype Content Model
Vision
Language and Cognition
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...