Related Experiment Video
Updated: Aug 21, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face.
Tong Shan1,2,3, Casper E Wenner4, Chenliang Xu5
1Department of Biomedical Engineering, 6927University of Rochester, Rochester, NY, USA.
Trends in Hearing
|November 17, 2022
Summary
Synthesized talking face videos significantly improve speech comprehension in noise. This deep neural network system offers a promising visual hearing aid, enhancing understanding even when audio is unclear.
Area of Science:
- Audiology and Speech Science
- Artificial Intelligence
- Human-Computer Interaction
Background:
- Speech comprehension is difficult in noisy environments.
- Visual cues from a talker's face significantly improve speech understanding.
- Deep neural networks (DNNs) can synthesize talking face videos from audio.
Purpose of the Study:
- To quantify the benefits of DNN-generated talking face videos for speech comprehension in noise.
- To compare synthesized audio-visual (AV) stimuli with natural AV stimuli and audio-only conditions.
Main Methods:
- Speech audio was masked at various signal-to-noise ratios (-9 to 0 dB).
- Subjects experienced three conditions: synthesized AV, natural AV, and audio-only.
- Speech comprehension was assessed by typed sentence recall and keyword recognition.
Main Results:
- Synthesized AV stimuli showed significant improvement over audio-only, falling between audio-only and natural AV performance.
- All subjects benefited from the synthesized AV condition.
- Performance improved with increasing signal-to-noise ratios across all conditions.
Conclusions:
- DNN-based synthesized talking face videos meaningfully enhance speech comprehension in noisy environments.
- This technology shows potential as a visual hearing aid.
- Further research can optimize DNN models for even greater comprehension benefits.
More Related Videos
Related Concept Videos
Facial Feedback Hypothesis
232
Charles Darwin proposed that facial expressions are an evolutionary adaptation for communication. He argued that these expressions are not influenced by culture but are universal across species. For example, a snarling expression with exposed teeth signals a threat in many animals, including humans. Darwin also suggested that displaying an emotion can intensify the feeling. Smiling, for example, could enhance one's sense of happiness. This idea laid the foundation for understanding the role...
232
Prosopagnosia
234
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
234
Nonconscious Mimicry
4.6K
Nonconscious mimicry occurs when individuals alter their mannerisms to match the behaviors and expressions of those nearby, without intention.
4.6K
Neural Circuits
1.4K
Neural circuits and neuronal pools are two of the main structures found in the nervous system. Neural circuits are networks of neurons that work together to carry out a specific task or process. They consist of interconnected neurons and glial cells, which provide structural and metabolic support.
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
1.4K

