Related Experiment Video
Updated: May 10, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Text-in-Image Enhanced Self-Supervised Alignment Model for Aspect-Based Multimodal Sentiment Analysis on Social
Xuefeng Zhao1, Yuxiang Wang1, Zhaoman Zhong1
1School of Computer Engineering, Jiangsu Ocean University, Lianyungang 222005, China.
This study introduces TESAM, a novel model for aspect-based multimodal sentiment analysis (ABMSA) that effectively analyzes social media content. TESAM improves sentiment analysis accuracy by incorporating text found within images.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Computer Vision
Background:
- Social media's growth necessitates advanced opinion mining and sentiment analysis on multimodal data.
- Aspect-based multimodal sentiment analysis (ABMSA) is crucial for fine-grained sentiment determination.
- Existing ABMSA methods struggle with social media data due to overlooked embedded image text.
Purpose of the Study:
- To develop a model that comprehensively analyzes multimodal information for ABMSA.
- To enhance sentiment analysis accuracy by integrating text extracted from images.
- To reduce noise interference and improve focus on relevant semantic features.
Main Methods:
- Proposed the Text-in-Image Enhanced Self-supervised Alignment Model (TESAM).
- Utilized Optical Character Recognition (OCR) to extract embedded text from images.
- Fused extracted text with visual features and incorporated aspect words for focused analysis.
- Employed self-supervised alignment pre-training to bridge the semantic gap between modalities using Euclidean distance and cosine similarity.
Main Results:
- TESAM demonstrated remarkable performance on three ABMSA benchmarks.
- The model effectively integrated visual and textual information, including embedded image text.
- Incorporating aspect words successfully reduced interference from irrelevant semantic features.
Conclusions:
- The proposed TESAM model significantly improves aspect-based multimodal sentiment analysis on social media data.
- Integrating embedded image text and aspect-guided feature selection enhances sentiment analysis accuracy.
- Self-supervised alignment is effective in mitigating the semantic gap between different modalities.
More Related Videos
06:19Integration of Animal Behavioral Assessment and Convolutional Neural Network to Study Wasabi-Alcohol Taste-Smell Interaction
Published on: August 16, 2024
07:12Protocol for Data Collection and Analysis Applied to Automated Facial Expression Analysis Technology and Temporal Analysis for Sensory Evaluation
Published on: August 26, 2016
Related Concept Videos
Social Proof
Stereotype Content Model
Social Facilitation
The Sense of Self: Reflected Self-Appraisal and Social Comparison
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Social Scripts