Related Experiment Video
Updated: Jul 11, 2026

Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
Multi-grained visual pivot-guided multi-modal neural machine translation with text-aware cross-modal contrastive
Junjun Guo1, Rui Su1, Junjie Ye2
1Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, Yunnan, 650500, China; Yunnan Key Laboratory of Artificial Intelligence, Kunming University of Science and Technology, Kunming, Yunnan, 650500, China.
This study introduces a novel multi-modal fusion strategy for neural machine translation, using visual information as a cross-lingual pivot to bridge semantic gaps between languages and improve translation quality.
Area of Science:
- Natural Language Processing
- Computer Vision
- Machine Translation
Background:
- Multi-modal neural machine translation (MNMT) aims to enhance text translation by integrating visual information.
- Semantic mismatch between image and text modalities poses a significant challenge in MNMT.
Purpose of the Study:
- To address the semantic mismatch problem in MNMT.
- To improve the performance of MNMT by leveraging visual information more effectively.
Main Methods:
- A multi-grained visual pivot-guided multi-modal fusion strategy is proposed.
- Cross-modal contrastive disentangling is employed to separate image information.
- Text-guided stacked cross-modal disentangling modules are introduced to disentangle images into MT-related and background visual information.
Main Results:
- The proposed approach significantly improves MNMT performance across four benchmark datasets.
- The method demonstrates superior results compared to existing state-of-the-art approaches.
- Analysis confirms the benefits of text-guided disentangling and visual pivot fusion.
Conclusions:
- The novel strategy effectively bridges linguistic gaps by using disentangled visual information as a cross-lingual pivot.
- This approach enhances cross-lingual alignment and boosts MNMT performance.
- The study contributes a promising direction for future research in multi-modal translation.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Multi-input and Multi-variable systems
In the absence of...

