Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Non-Verbal Cues01:29

Non-Verbal Cues

286
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
286
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

942
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
942
Labeling Emotion01:20

Labeling Emotion

604
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
604
Emotional Expression01:26

Emotional Expression

956
Emotional expression encompasses how individuals convey their emotions through verbal communication and non-verbal cues. These non-verbal actions include facial expressions, body language, and physical gestures, such as frowning or smiling. Among these, facial expressions play a crucial role in emotional expression and are understood universally, indicating a biological basis for how humans communicate emotions.
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
956
Coping Strategies: Emotion Focused01:20

Coping Strategies: Emotion Focused

443
Emotion-focused coping refers to a set of strategies aimed at managing the emotional impact of stressors, rather than directly addressing their causes. This approach involves altering one's emotional response to stressful situations to reduce their psychological effects. For example, individuals might talk with a friend or engage in activities like journaling to express their feelings. Such actions can help achieve emotional clarity or release, providing the psychological stability needed...
443
Auditory Perception01:17

Auditory Perception

1.0K
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the...
1.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Reduced Dose-Direct Oral Anticoagulant Vs Dual Antiplatelet Therapy After Left Atrial Appendage Closure in Patients With Nonvalvular Atrial Fibrillation: A Systematic Review and Meta Analysis.

The Annals of pharmacotherapy·2026
Same author

Efficacy and safety of optical coherence tomography-guided versus angiography-guided percutaneous coronary intervention: a systematic review and meta-analysis.

Annals of medicine and surgery (2012)·2026
Same author

ATEdrug: A reliable human-in-the-loop annotation scheme for aspect term extraction and polarity detection in drug reviews.

PloS one·2026
Same author

Next-generation digital twin model with unobtrusive RF multi-sensing for AI-based human monitoring.

Scientific reports·2026
Same author

Deep learning analysis for enhanced prediction of heat transfer in Maxwell hybrid nanofluids with non-Fourier law and radiation effects.

Scientific reports·2026
Same author

A novel deep semantic- and vision-based self-attention architecture for skin cancer classification.

Digital health·2026

Related Experiment Video

Updated: Jan 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K

Audio-visual speech enhancement in noisy environments using emotion-based contextual cues.

Tassadaq Hussain1, Nasir Saleem1, Kia Dashtipour1

  • 1School of Computing Engineering and the Built Environment, Edinburgh Napier University, Edinburgh, EH105DT, United Kingdom.

The Journal of the Acoustical Society of America
|January 16, 2026
PubMed
Summary

This study introduces an emotion-aware audio-visual speech enhancement (EAVSE) system. Incorporating emotional cues from facial landmarks significantly improves speech clarity in noisy conditions.

More Related Videos

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.4K
Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
11:39

Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique

Published on: September 7, 2022

2.6K

Related Experiment Videos

Last Updated: Jan 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.4K
Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
11:39

Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique

Published on: September 7, 2022

2.6K

Area of Science:

  • Signal Processing
  • Artificial Intelligence
  • Human-Computer Interaction

Background:

  • Background noise severely impacts speech intelligibility in real-world scenarios.
  • Current audio-visual speech enhancement (AVSE) methods struggle with dynamic, noisy environments.
  • Emotional context is an underutilized feature in speech enhancement.

Purpose of the Study:

  • To propose and evaluate a novel emotion-aware audio-visual speech enhancement (EAVSE) system.
  • To investigate the impact of incorporating emotional features from facial landmarks into AVSE.
  • To enhance speech clarity and intelligibility in challenging acoustic conditions.

Main Methods:

  • Developed a deep learning-based EAVSE system utilizing auditory, visual, and emotional information.
  • Extracted emotional features from facial landmarks and fused them with audio-visual data.
  • Employed a UNet-based encoder-decoder network for joint multi-modal learning.
  • Utilized a distortion-inspired loss function, specifically scale-invariant signal-to-distortion ratio (S-SNR), for optimization.
  • Trained and evaluated the model on the Carnegie Mellon University Multimodal Opinion Sentiment and Emotion Intensity dataset.

Main Results:

  • The EAVSE system achieved significant improvements in objective metrics (PESQ, STOI) and subjective speech quality.
  • Demonstrated superior performance compared to benchmark AVSE and audio-only systems, especially in low SNR environments (≤ -7.5 dB).
  • Achieved Δ STOI of 7.32%, Δ PESQ of 0.33, and Δ S-SNR of 7.8 dB over noisy benchmarks at 0 dB SNR.

Conclusions:

  • Emotional contextual cues derived from facial landmarks are effective for improving audio-visual speech enhancement.
  • The proposed EAVSE system offers a robust solution for speech enhancement in noisy and dynamic environments.
  • The findings highlight the potential of multi-modal deep learning integrating emotional understanding for better human-computer interaction.