Related Experiment Video
Updated: Nov 1, 2025

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement
1National Engineering Laboratory for Speech and Language Information Processing, University of Science and Technology of China, Hefei, Anhui, China.
This study introduces a visual embedding method to enhance speech by synchronizing lip movements. The multi-modal approach significantly improves speech quality and intelligibility compared to audio-only or visual-only systems.
Area of Science:
- Speech processing
- Computer vision
- Machine learning
Background:
- Speech enhancement aims to improve audio quality in noisy environments.
- Existing methods often rely solely on audio features.
- Visual cues from lip movements offer complementary information for speech enhancement.
Purpose of the Study:
- To propose a novel visual embedding approach for embedding aware speech enhancement (EASE).
- To synchronize visual lip frames at phone and articulation place levels for improved speech enhancement.
- To develop a multi-modal EASE (MEASE) framework integrating audio and visual features.
Main Methods:
- Extracting visual embeddings from lip frames using pre-trained phone or articulation place recognizers for visual-only EASE (VEASE).
- Developing a MEASE framework by extracting audio-visual embeddings through information intersection of audio and visual features.
- Utilizing subword-based embeddings and focusing on articulation place level for enhanced visual feature extraction.
Main Results:
- Subword-based VEASE outperforms conventional word-level embedding for speech enhancement.
- Visual embedding at the articulation place level shows superior performance compared to the phone level.
- The proposed MEASE framework significantly enhances speech quality and intelligibility over single-modality systems.
Conclusions:
- Visual embedding synchronization at phone and articulation place levels effectively improves speech enhancement.
- Multi-modal EASE (MEASE) leveraging audio-visual feature complementarity offers substantial gains in speech quality and intelligibility.
- The proposed approach represents a significant advancement in robust speech enhancement technology.
More Related Videos
Related Concept Videos
Facial Feedback Hypothesis
Non-Verbal Cues
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Chunking and Rehearsal in Sensory Memory
Auditory Perception
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.

