Related Experiment Video
Updated: Jun 13, 2025

05:51
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
9.0K
MelTrans: Mel-Spectrogram Relationship-Learning for Speech Emotion Recognition via Transformers.
Hui Li1,2, Jiawen Li2, Hai Liu2
1School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China.
Sensors (Basel, Switzerland)
|September 14, 2024
Summary
This study introduces MelTrans, a Transformer-based model for speech emotion recognition (SER). MelTrans effectively captures subtle emotional cues and long-range dependencies in speech, outperforming previous benchmarks on key datasets.
Area of Science:
- Artificial Intelligence
- Human-Computer Interaction
- Signal Processing
Background:
- Speech emotion recognition (SER) is crucial for natural human-computer interaction.
- Existing SER methods struggle with subtle emotions and noisy environments.
- Advanced feature extraction and dependency modeling are needed for robust SER.
Purpose of the Study:
- To introduce MelTrans, a novel Transformer-based model for enhanced speech emotion recognition.
- To address challenges in detecting subtle emotional nuances and recognizing emotions in noisy speech.
- To improve the accuracy and robustness of SER systems.
Main Methods:
- Developed MelTrans, a dual-stream Transformer-based model.
- Utilized speech mel-spectrograms to capture broad dependencies.
- Focused on learning core features and long-range dependencies within speech data.
Main Results:
- MelTrans achieved 92.52% accuracy on the EmoDB dataset.
- MelTrans achieved 76.54% accuracy on the IEMOCAP dataset.
- Demonstrated superior performance in capturing critical cues and long-range dependencies.
Conclusions:
- MelTrans effectively addresses complex challenges in speech emotion recognition.
- The model sets new benchmarks for SER on the EmoDB and IEMOCAP datasets.
- Highlights the potential of Transformer architectures for nuanced emotion detection in speech.

