Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types Of Transformers01:16

Types Of Transformers

943
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
943
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
93
Association Areas of the Cortex01:21

Association Areas of the Cortex

4.9K
Association areas are regions of the cerebral cortex that do not have a specific sensory or motor function. Instead, they integrate and interpret information from various sources to enable higher cognitive processes such as memory, learning, and decision-making. Some key association areas include the following:
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
4.9K
The Ideal Transformer01:26

The Ideal Transformer

343
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
343
Facial Feedback Hypothesis01:24

Facial Feedback Hypothesis

117
Charles Darwin proposed that facial expressions are an evolutionary adaptation for communication. He argued that these expressions are not influenced by culture but are universal across species. For example, a snarling expression with exposed teeth signals a threat in many animals, including humans. Darwin also suggested that displaying an emotion can intensify the feeling. Smiling, for example, could enhance one's sense of happiness. This idea laid the foundation for understanding the role...
117
Cognitive Theories: Schachter-Singer Theory of Emotion01:20

Cognitive Theories: Schachter-Singer Theory of Emotion

228
Stanley Schachter and Jerome Singer proposed the two-factor theory of emotion, which emphasizes the interplay between physiological arousal and cognitive labeling in forming emotional experiences. This theory suggests that emotions are not simply a result of physiological responses but rather a combination of these responses and the individual's cognitive interpretation of them.
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...
228

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Corrigendum to "GinDB-AI: An integrated database of Panax-derived compounds and an AI-driven platform for multidimensional information and biological activity prediction" [J Ginseng Res 50/3 (2026) 100986].

Journal of ginseng research·2026
Same author

CONTRA-IL6: an interpretable hybrid convolutional neural network and Transformer framework for accurate prediction of interleukin-6-inducing peptides using protein language models.

Briefings in bioinformatics·2026
Same author

GinDB-AI: An integrated ginsenoside database and AI-driven platform for multidimensional information and biological activity prediction.

Journal of ginseng research·2026
Same author

Mulaqua: An interpretable multimodal deep learning framework for identifying PMT/vPvM substances in drinking water.

Journal of hazardous materials·2025
Same author

A deep learning-based framework for standardized analysis of trabecular bone compartments from micro-CT imaging data in the mouse tibia.

Scientific reports·2025
Same author

HyPepTox-Fuse: An interpretable hybrid framework for accurate peptide toxicity prediction fusing protein language model-based embeddings with conventional descriptors.

Journal of pharmaceutical analysis·2025

Related Experiment Video

Updated: May 28, 2025

Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

19.9K

MemoCMT: multimodal emotion recognition using cross-modal transformer-based feature fusion.

Mustaqeem Khan1, Phuong-Nam Tran2, Nhat Truong Pham3

  • 1College of Information Technology, United Arab Emirates University (UAEU), Al Ain, Abu Dhabi, United Arab Emirates, 5551, Al Ain, UAE.

Scientific Reports
|February 14, 2025
PubMed
Summary

MemoCMT enhances speech emotion recognition by integrating audio and text features using a novel cross-modal transformer. This efficient system achieves high accuracy on benchmark datasets, showing promise for real-world applications.

Keywords:
Cross-modal transformerDeep learningFeature fusionMultimodal emotion recognitionSpeech emotion recognition

More Related Videos

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
07:13

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities

Published on: October 27, 2023

1.0K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

8.9K

Related Experiment Videos

Last Updated: May 28, 2025

Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

19.9K
Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
07:13

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities

Published on: October 27, 2023

1.0K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

8.9K

Area of Science:

  • Artificial Intelligence
  • Speech Processing
  • Machine Learning

Background:

  • Transformer models excel in speech emotion recognition but are computationally expensive.
  • Convolutional neural networks are faster but limited in capturing long-range speech patterns.
  • Existing methods face challenges balancing accuracy and computational efficiency.

Purpose of the Study:

  • To develop an efficient and accurate speech emotion recognition system.
  • To introduce a novel cross-modal transformer (CMT) for analyzing local and global speech features with text.
  • To leverage pre-trained models like HuBERT and BERT for feature extraction.

Main Methods:

  • Proposed MemoCMT system utilizing a cross-modal transformer (CMT).
  • Integration of HuBERT for audio feature extraction and BERT for text analysis.
  • Application of various fusion techniques post-feature integration for emotion classification.

Main Results:

  • MemoCMT achieved high performance on IEMOCAP and ESD benchmark corpora.
  • The CMT with min aggregation yielded the highest unweighted accuracy (UW-Acc) of 81.33% and 91.93%.
  • Weighted accuracy (W-Acc) reached 81.85% and 91.84% respectively.

Conclusions:

  • MemoCMT demonstrates strong generalization capacity and robustness.
  • The system offers an efficient solution for speech emotion recognition.
  • Publicly available implementation promotes reproducibility and further research.