Related Experiment Video
Updated: Jan 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Audio-visual speech enhancement in noisy environments using emotion-based contextual cues
Tassadaq Hussain1, Nasir Saleem1, Kia Dashtipour1
1School of Computing Engineering and the Built Environment, Edinburgh Napier University, Edinburgh, EH105DT, United Kingdom.
Abstract:
In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Existing audio-visual speech enhancement (AVSE) techniques often pose challenges in dynamic and noisy conditions. This study examines the inclusion of emotional features as a novel contextual cue within the AVSE framework. We analyze that incorporating emotional understanding from facial landmarks improves speech enhancement performance. We propose a deep learning-based emotion-aware audio-visual speech enhancement system (EAVSE) that uses auditory, visual, and emotional information. The proposed EAVSE extracts emotional features from facial landmarks and combines them with audio and visual modalities. Enriched multi-model data are processed by a UNet-based encoder-decoder network for joint learning and optimization. The network iteratively refines the feature representation, guided by a distortion-inspired loss function. We train and evaluate the model on the Carnegie Mellon University Multimodal Opinion Sentiment and Emotion Intensity dataset, known for its diverse audio-visual recordings with annotated emotions. Compared to AVSE benchmark and audio-only speech enhancement systems, the proposed model achieves significant improvements in both objective [Perceptual Evaluation of Speech Quality (PESQ), Short-Time Objective Intelligibility (STOI)] and subjective speech quality metrics. In particular, the scale-invariant signal-to-distortion ratio loss function demonstrates superior performance. This suggests the usefulness of the emotional contextual cues for AVSE. The experimental findings demonstrate the effectiveness of the AVSE, particularly in challenging noisy environments [signal-to-noise ratio (SNR) ≤ -7.5 dB]. The proposed model achieved Δ STOI of 7.32%, Δ PESQ of 0.33, and Δ S-SNR of 7.8 dB over noisy benchmark at 0 dB SNR.
More Related Videos
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
11:39Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
Published on: September 7, 2022
Related Concept Videos
Non-Verbal Cues
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Labeling Emotion
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Coping Strategies: Emotion Focused
Auditory Perception