Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

A text-guided cross-hierarchical fusion and multi-task learning framework for multimodal sentiment analysis.

Minghui Zhu1, Yushan Pan2, Nan Xiang2

  • 1Department of Computing, Xi'an Jiaotong-Liverpool University, Suzhou, 215123, Jiangsu, China; School of Computer Science and Informatics, The University of Liverpool, Liverpool, L69 7ZX, United Kingdom.

Neural Networks : the Official Journal of the International Neural Network Society
|April 22, 2026
PubMed
Summary

Related Concept Videos

Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

539
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
539

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Electricity generation from food wastes and microbial community structure in microbial fuel cells.

Bioresource technology·2013
Same author

Plasmon-enhanced photothermoelectric conversion in chemical vapor deposited graphene p-n junctions.

Journal of the American Chemical Society·2013
Same author

Rapid and sensitive detection of vinorelbine in the urine of tumor patients by capillary electrophoresis with tris(2,2'-bipyridyl)ruthenium(II)-based electrochemiluminescence assay.

Analytical sciences : the international journal of the Japan Society for Analytical Chemistry·2013
Same author

Synthesis and biological evaluation of 4-phenoxy-6,7-disubstituted quinolines possessing semicarbazone scaffolds as selective c-Met inhibitors.

Archiv der Pharmazie·2013
Same author

The effects of external electric field: creating non-zero first hyperpolarizability for centrosymmetric benzene and strongly enhancing first hyperpolarizability for non-centrosymmetric edge-modified graphene ribbon H2N-(3,3)ZGNR-NO2.

Journal of molecular modeling·2013
Same author

Msx2 plays a critical role in lens epithelium cell cycle control.

International journal of ophthalmology·2013

This study introduces a novel multi-task learning framework for multimodal sentiment analysis (MSA). The proposed method enhances feature utilization and jointly models hierarchical and modality-specific characteristics, improving sentiment prediction accuracy.

Area of Science:

  • Artificial Intelligence
  • Natural Language Processing
  • Computer Vision

Background:

  • Existing multimodal sentiment analysis (MSA) methods struggle with insufficient textual information utilization, limited hierarchical feature modeling, and inadequate exploration of modality-specific characteristics.
  • These limitations can lead to overlooking emotional patterns and hinder effective cross-hierarchical and modality-specific feature learning.

Purpose of the Study:

  • To propose a novel multi-task learning framework to address the challenges in existing multimodal sentiment analysis (MSA) methods.
  • To improve the utilization of textual modality information and jointly model hierarchical multimodal features.
  • To enhance the exploration of independent characteristics of each modality for more effective sentiment analysis.

Main Methods:

Keywords:
Hierarchical feature fusionMultimodal sentiment analysisTextual information enhancementUnimodal characteristic extraction

Related Experiment Videos

  • A multi-task learning framework is proposed, jointly modeling three unimodal and one multimodal sentiment prediction tasks.
  • Three innovative modules are introduced: Raw Cross-Modal Information Fusion (RCMIF) using graph convolution for shallow features, Cleaned Cross-Modal Information Fusion (CCMIF) with dynamic graph convolution and attention for deep features, and Bilinear Attention Deep Independent Characteristics Mining (BADIC) for unimodal characteristics.
  • Textual guidance is exploited in BADIC and CCMIF, while RCMIF and CCMIF collaborate to learn cross-hierarchical features.

Main Results:

  • The proposed framework demonstrates improved performance on multiple benchmark datasets (MOSI, MOSEI, CH-SIMS, CH-SIMS-V2).
  • Achieved approximately 0.5%-1% improvement in binary classification accuracy and F1 score compared to existing state-of-the-art methods.
  • The framework effectively optimizes network parameters and promotes learning of modality-specific characteristics through unimodal branches.

Conclusions:

  • The proposed multi-task learning framework effectively addresses key challenges in multimodal sentiment analysis.
  • The novel fusion and characteristic mining modules enhance the learning of both cross-hierarchical and modality-specific features.
  • The framework offers a promising direction for advancing the performance and robustness of multimodal sentiment analysis systems.