Related Experiment Video
Updated: Sep 10, 2025

Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
C3aptioner: Improving change captioning by leveraging momentum cross-view and cross-modality contrastive learning
Lin Deng1, Borui Kang2, Yuzhong Zhong3
1College of Electrical Engineering, Sichuan University, Chengdu, 610065, PR China; National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University, Chengdu, 610064, PR China.
This study introduces C³aptioner, a novel change captioning model. It improves descriptions by identifying visual changes and retaining unchanged elements using dual contrastive learning.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Change captioning aims to describe visual differences between images.
- Existing methods focus on detecting and describing changes, neglecting persistent elements and vision-language links.
- Effective change captioning requires describing both changes and unchanged semantic elements, linking visual cues to language.
Purpose of the Study:
- To develop an advanced change captioning model, C³aptioner, that addresses limitations of prior work.
- To enhance descriptions by incorporating both changed and unchanged visual information.
- To establish a robust connection between visual features and linguistic descriptions.
Main Methods:
- The C³aptioner model utilizes dual momentum contrastive learning objectives.
- It employs intra-image and inter-image Transformer encoders for visual feature extraction.
- Cross-view and cross-modality contrastive learning objectives align visual and textual representations, addressing viewpoint variations and modality gaps.
Main Results:
- The C³aptioner model achieves state-of-the-art performance across five datasets.
- Significant improvements were observed in challenging scenarios with extreme viewpoint changes.
- The dual contrastive approach effectively models both changed and unchanged elements, enhancing vision-language correspondence.
Conclusions:
- C³aptioner provides more contextually rich and human-like change descriptions.
- The proposed dual contrastive learning strategy is effective for change captioning.
- The model demonstrates superior performance, particularly in handling viewpoint variations.
More Related Videos
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Improving Translational Accuracy
Observational Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Learning Disabilities
Dyslexia
Dyslexia is a...
Associative Learning
Classical conditioning, also known...
The Anchoring-and-Adjustment Heuristic