Related Experiment Video
Updated: Nov 26, 2025

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
20.3K
Cross-Domain Image Captioning via Cross-Modal Retrieval and Model Adaptation
Summary
This study introduces a novel cross-modal retrieval method to improve cross-domain image captioning. The approach generates pseudo image-sentence pairs for adapting models to new domains with unpaired data, achieving strong performance.
Area of Science:
- Computer Vision
- Natural Language Processing
- Machine Learning
Background:
- Large-scale paired image-sentence datasets fuel image captioning success.
- Collecting domain-specific paired data is resource-intensive.
- Transfer learning offers a solution for adapting models to new domains with limited data.
Purpose of the Study:
- To develop an effective method for cross-domain image captioning using unpaired data.
- To leverage cross-modal retrieval for generating pseudo image-sentence pairs in target domains.
- To adapt image captioning models trained on source domains to new target domains.
Main Methods:
- Propose a cross-modal retrieval aided approach for cross-domain image captioning.
- Utilize an iterative cross-modal retrieval process to generate and refine pseudo image-sentence pairs.
- Employ an adaptive image captioning model with self-attention, fine-tuned on pseudo pairs.
Main Results:
- The proposed method achieves competitive or superior performance compared to state-of-the-art approaches across multiple target domains.
- Demonstrated effectiveness in cross-domain image captioning tasks using MSCOCO as the source domain.
- Successfully extended to cross-domain video captioning, validating the method's versatility.
Conclusions:
- The cross-modal retrieval approach effectively addresses the challenge of limited paired data in cross-domain image captioning.
- Iterative refinement of pseudo pairs enhances model adaptation and captioning quality.
- The method shows promise for various cross-modal generation tasks beyond image captioning.
More Related Videos
Related Concept Videos
Retrieval
282
Retrieval is the process of getting information out of memory storage and back into conscious awareness. This ability is essential for daily tasks like brushing hair and teeth, driving to work, and performing job duties. Retrieval occurs in three ways: recall, recognition, and relearning.
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
282
Stereotype Content Model
15.1K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.1K
Cross-reactivity
32.2K
Overview
32.2K

