Related Experiment Videos
CLAMP: Contrastive learning with adaptive multi-loss and progressive fusion for multimodal aspect-based sentiment
Xiaoqiang He1, Qiumei Pu1, Xin Luo1
1Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing, 100081, China; School of Information Engineering, Minzu University of China, Beijing, 100081, China.
Summary
This study introduces CLAMP, a novel framework for multimodal aspect-based sentiment analysis (MABSA). CLAMP improves cross-modal alignment and representation consistency for more accurate sentiment analysis in image-text data.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Computer Vision
Background:
- Multimodal aspect-based sentiment analysis (MABSA) is crucial for applications like product reviews.
- Current MABSA methods struggle with cross-modal alignment noise and inconsistent representations.
- Existing approaches often fail to connect textual aspect terms with relevant visual information.
Purpose of the Study:
- To develop an end-to-end framework, CLAMP, to address limitations in MABSA.
- To enhance fine-grained alignment and consistency in cross-modal representations.
- To improve the accuracy of sentiment analysis in paired image-text data.
Main Methods:
- Introduced CLAMP, a Contrastive Learning framework with Adaptive Multi-loss and Progressive Attention Fusion.
- Implemented a Progressive Attention Fusion network for hierarchical cross-modal interaction and noise suppression.
- Utilized multi-task contrastive learning for global and local alignment, and adaptive multi-loss aggregation for uncertainty-based loss calibration.
Main Results:
- CLAMP demonstrates superior performance on standard benchmarks.
- The framework effectively suppresses irrelevant visual noise through progressive attention fusion.
- Achieved enhanced cross-modal representation consistency via multi-task contrastive learning.
Conclusions:
- CLAMP offers a significant advancement in multimodal aspect-based sentiment analysis.
- The proposed modules effectively tackle challenges in cross-modal alignment and representation.
- The framework outperforms existing state-of-the-art methods in MABSA tasks.