Related Experiment Video
Updated: May 6, 2026

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
373
Selective Acoustic Feature Enhancement for Speech Emotion Recognition With Noisy Speech
Seong-Gyun Leem1, Daniel Fulford2, Jukka-Pekka Onnela3
1Department of Electrical and Computer Engineering, University of Texas at Dallas, Richardson, TX 75080 USA.
Summary
This study introduces a novel speech enhancement method for speech emotion recognition (SER) systems. By selectively enhancing only weak acoustic features, the proposed approach significantly improves emotion recognition performance in noisy conditions.
Area of Science:
- Speech processing
- Machine learning
- Acoustic analysis
Background:
- Real-world speech emotion recognition (SER) systems face challenges with background noise.
- Speech enhancement (SE) modules can improve speech quality but may degrade crucial SER features.
- Existing SE methods risk altering robust acoustic features essential for accurate emotion recognition.
Purpose of the Study:
- To develop a targeted speech enhancement strategy for SER systems operating in noisy environments.
- To enhance only the weak acoustic features that negatively impact emotion recognition performance.
- To preserve robust features that are resilient to environmental variations.
Main Methods:
- Identified weak features using multiple single-feature acoustic models trained on clean speech.
- Ranked features based on performance, robustness, and a combined joint rank.
- Selectively enhanced identified weak low-level descriptors (LLDs), preserving robust features.
Main Results:
- Directly enhancing weak LLDs outperformed extracting LLDs from fully enhanced speech.
- Achieved significant performance gains: 17.7% (arousal), 21.2% (dominance), and 3.3% (valence) at 10dB SNR.
- Outperformed a system that enhanced all LLDs on the MSP-Podcast corpus.
Conclusions:
- Targeted enhancement of weak features is a more effective strategy for SER in noisy conditions.
- The proposed method preserves discriminative acoustic information crucial for robust emotion recognition.
- This approach offers substantial improvements in SER accuracy across various emotional dimensions.
More Related Videos
Related Concept Videos
Labeling Emotion
1.0K
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
1.0K
Non-Verbal Cues
784
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
784

