Related Experiment Video
Updated: Jan 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Audio-visual speech enhancement in noisy environments using emotion-based contextual cues
Tassadaq Hussain1, Nasir Saleem1, Kia Dashtipour1
1School of Computing Engineering and the Built Environment, Edinburgh Napier University, Edinburgh, EH105DT, United Kingdom.
This study introduces an emotion-aware audio-visual speech enhancement (EAVSE) system. Incorporating emotional cues from facial landmarks significantly improves speech clarity in noisy conditions.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Human-Computer Interaction
Background:
- Background noise severely impacts speech intelligibility in real-world scenarios.
- Current audio-visual speech enhancement (AVSE) methods struggle with dynamic, noisy environments.
- Emotional context is an underutilized feature in speech enhancement.
Purpose of the Study:
- To propose and evaluate a novel emotion-aware audio-visual speech enhancement (EAVSE) system.
- To investigate the impact of incorporating emotional features from facial landmarks into AVSE.
- To enhance speech clarity and intelligibility in challenging acoustic conditions.
Main Methods:
- Developed a deep learning-based EAVSE system utilizing auditory, visual, and emotional information.
- Extracted emotional features from facial landmarks and fused them with audio-visual data.
- Employed a UNet-based encoder-decoder network for joint multi-modal learning.
- Utilized a distortion-inspired loss function, specifically scale-invariant signal-to-distortion ratio (S-SNR), for optimization.
- Trained and evaluated the model on the Carnegie Mellon University Multimodal Opinion Sentiment and Emotion Intensity dataset.
Main Results:
- The EAVSE system achieved significant improvements in objective metrics (PESQ, STOI) and subjective speech quality.
- Demonstrated superior performance compared to benchmark AVSE and audio-only systems, especially in low SNR environments (≤ -7.5 dB).
- Achieved Δ STOI of 7.32%, Δ PESQ of 0.33, and Δ S-SNR of 7.8 dB over noisy benchmarks at 0 dB SNR.
Conclusions:
- Emotional contextual cues derived from facial landmarks are effective for improving audio-visual speech enhancement.
- The proposed EAVSE system offers a robust solution for speech enhancement in noisy and dynamic environments.
- The findings highlight the potential of multi-modal deep learning integrating emotional understanding for better human-computer interaction.
More Related Videos
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
11:39Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
Published on: September 7, 2022
Related Concept Videos
Non-Verbal Cues
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Labeling Emotion
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Coping Strategies: Emotion Focused
Auditory Perception