Related Experiment Video
Updated: Sep 10, 2025

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
Published on: March 8, 2024
Aspect-level multimodal sentiment analysis model based on multi-scale feature extraction
1Beijing Institute of Graphic Communication, Beijing, 102600, China.
Abstract:
In existing multimodal sentiment analysis methods, only the last layer output of BERT is typically used for feature extraction, neglecting abundant information from intermediate layers. This paper proposes an Aspect-level Multimodal Sentiment Analysis Model with Multi-scale Feature Extraction (AMSAM-MFE). The model conducts sentiment analysis on both text and images. For text feature extraction, it incorporates a Multi-scale Layer module based on BERT and utilizes aspect terms to supervise text feature extraction, enhancing text processing performance. For image feature extraction, the model employs a pre-trained Resnest269 model with a specially designed Supervision Layer to improve effectiveness. For feature fusion, the Tensor Fusion Network method is adopted to achieve comprehensive interaction between visual and textual features. Experimental comparisons with other multimodal sentiment analysis models on Twitter2015 and Twitter2017 datasets demonstrated that the proposed multi-scale feature extraction model achieved improved accuracy and F1 scores in aspect-level multimodal sentiment analysis tasks, showing superior classification effectiveness compared to traditional multimodal sentiment analysis models.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Upsampling
Scaling
Root Mean Square
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...

