Related Experiment Video
Updated: Apr 18, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.1K
Deep cross-modal affective memory networks with adaptive multi-source heterogeneous transfer learning in speech
Xiaofen Zhao1, Jingchao Liu2, Lei Lin3
1School of Computer Science, Xijing University, Xi'an, 710123, China. 1418440678@qq.com.
Scientific Reports
|April 16, 2026
Summary
This study introduces novel Deep Cross-Modal Emotional Memory Networks (DCM-EMNet) and Adaptive Multi-source Heterogeneous Transfer Learning Frameworks (AMS-HTLF) to significantly improve speech emotion recognition accuracy and model generalization.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Natural Language Processing
Background:
- Speech emotion recognition (SER) faces challenges in multimodal data fusion and cross-domain generalization.
- Existing methods struggle with integrating diverse data sources and migrating knowledge effectively.
Purpose of the Study:
- To propose an innovative Deep Cross-Modal Emotional Memory Network (DCM-EMNet) for enhanced SER.
- To develop an Adaptive Multi-source Heterogeneous Transfer Learning Framework (AMS-HTLF) for robust cross-domain SER.
Main Methods:
- DCM-EMNet utilizes multi-level feature fusion, a dynamic affective memory mechanism, and cross-modal consistency constraints for bimodal speech and text data.
- AMS-HTLF implements adaptive feature alignment and heterogeneous label mapping for effective data migration across domains.
Main Results:
- The proposed methods significantly improve accuracy in speech emotion recognition across multiple datasets.
- Experimental results validate the effectiveness and practicality of DCM-EMNet and AMS-HTLF.
- Enhanced generalization ability of the model on target domains was observed.
Conclusions:
- The study offers new research perspectives and methods for speech emotion recognition.
- The work expands innovative ideas in multimodal learning and transfer learning.
- The proposed frameworks effectively address multimodal data fusion and cross-domain SER challenges.
