Related Experiment Videos
Perception of visible speech: influence of spatial quantization
1Department of Psychology, University of California at Santa Cruz 95064, USA.
Perception
|January 1, 1997
Summary
Speech reading relies on specific visual features for understanding spoken language. A fuzzy-logical model accurately describes how these visual cues are processed, similar to facial recognition.
Area of Science:
- Auditory and Speech Science
- Cognitive Psychology
- Computer Vision
Background:
- Visible speech reading, or lip-reading, is a complex perceptual task.
- Understanding the key visual features used in speech reading is crucial for developing assistive technologies and improving human-computer interaction.
- Previous models of pattern recognition have not fully explained the nuances of speech perception from visual information alone.
Purpose of the Study:
- To identify the functional visual features critical for distinguishing between different speech sounds.
- To evaluate the efficacy of various pattern recognition models in explaining speech reading performance.
- To determine if speech reading follows similar perceptual principles as other visual recognition tasks.
Main Methods:
- Nine consonant-vowel syllables were presented visually under varying degrees of spatial degradation (quantization).
- Listener performance was measured across different degradation levels.
- Confusion matrices generated from performance data were analyzed using computational models, including a fuzzy-logical model and an additive model.
Main Results:
- Speech reading performance declined with increased spatial quantization but remained robust at moderate degradation levels.
- Six specific visual features were identified as essential for discriminating between the tested consonant sounds.
- The fuzzy-logical model significantly outperformed the additive model in describing the observed confusion patterns.
Conclusions:
- Visible speech reading involves the extraction and integration of specific functional visual features.
- The fuzzy-logical model provides a superior framework for understanding the cognitive processes underlying speech reading.
- Speech reading shares fundamental pattern recognition mechanisms with other visual perception domains, such as facial recognition and affect perception.