Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types Of Transformers01:16

Types Of Transformers

1.0K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.0K
Labeling Emotion01:20

Labeling Emotion

189
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
189
The Ideal Transformer01:26

The Ideal Transformer

438
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
438
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

134
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
134
Transformers01:26

Transformers

1.1K
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.1K
Association Areas of the Cortex01:21

Association Areas of the Cortex

5.6K
Association areas are regions of the cerebral cortex that do not have a specific sensory or motor function. Instead, they integrate and interpret information from various sources to enable higher cognitive processes such as memory, learning, and decision-making. Some key association areas include the following:
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
5.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Intracellular structural modifications of natural peptidoglycan fragments preceding NOD2 signaling in mammalian cells.

Proceedings of the National Academy of Sciences of the United States of America·2026
Same author

Serum Response Factor Regulates CCN1 to Exacerbate Acute Kidney Injury Through Facilitation of Ferroptosis-Related Injury.

Nephrology (Carlton, Vic.)·2026
Same author

The potential mechanisms of exercise-regulated mechanically sensitive ion channels in promoting spinal cord injury repair: a hypothesis-driven narrative review.

Reviews in the neurosciences·2026
Same author

Reduced Indocyanine Green Clearance Is Associated with Enteral Feeding Intolerance in Septic Patients Without Overt Liver Injury.

Journal of clinical medicine·2026
Same author

LabSage: Structural-Semantic Decoupling for Enhanced Retrieval-Augmented Generation in Clinical Laboratories.

AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science·2026
Same author

A multimodal generative model for structured and unstructured electronic health records.

npj health systems·2026

Related Experiment Video

Updated: Jul 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

Multimodal transformer augmented fusion for speech emotion recognition.

Yuanyuan Wang1, Yu Gu1, Yifei Yin2

  • 1School of Artificial Intelligence, Xidian University, Xi'an, China.

Frontiers in Neurorobotics
|June 7, 2023
PubMed
Summary

This study introduces a novel multimodal transformer augmented fusion method for speech emotion recognition. The approach enhances multimodal emotional representation by effectively integrating diverse data, outperforming existing methods.

Keywords:
hybrid fusionmodal interactionmultimodal enhancementspeech emotion recognitiontransformer

More Related Videos

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.1K
Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
09:44

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology

Published on: March 8, 2024

4.9K

Related Experiment Videos

Last Updated: Jul 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.1K
Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
09:44

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology

Published on: March 8, 2024

4.9K

Area of Science:

  • Artificial Intelligence
  • Human-Computer Interaction
  • Signal Processing

Background:

  • Speech emotion recognition (SER) is complex due to subjective and ambiguous emotional expression.
  • Multimodal SER methods show promise but struggle with integrating heterogeneous data.
  • Existing fusion techniques often overlook fine-grained interactions between modalities.

Purpose of the Study:

  • To develop an advanced fusion strategy for improved multimodal speech emotion recognition.
  • To address the challenge of effectively integrating information from different modalities.
  • To capture fine-grained interactions within and between speech and text modalities.

Main Methods:

  • Proposed a multimodal transformer augmented fusion method utilizing a hybrid fusion strategy.
  • Implemented a Model-fusion module with Cross-Transformer Encoders for multimodal emotional representation.
  • Employed feature-level and model-level fusion for enhanced speech feature extraction using text and multimodal features.

Main Results:

  • The proposed method demonstrated superior performance compared to state-of-the-art approaches.
  • Achieved significant improvements on the IEMOCAP and MELD datasets.
  • Successfully captured fine-grained modal interactions for more accurate emotion recognition.

Conclusions:

  • The multimodal transformer augmented fusion method offers a breakthrough in integrating heterogeneous multimodal data for SER.
  • The hybrid fusion strategy and Cross-Transformer Encoders effectively enhance multimodal emotional representation.
  • This approach advances the field of speech emotion recognition by enabling more nuanced understanding of emotional cues.