Related Experiment Video
Updated: Jul 7, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Lipreading from color video
Summary
This study presents a lipreading system using only lip color video for word recognition. The visual-only system achieved 94% accuracy for ten isolated words.
Area of Science:
- Computer Vision
- Biomedical Engineering
- Speech Recognition
Background:
- Traditional speech recognition relies heavily on acoustic data, limiting its use in noisy environments or for individuals with speech impairments.
- Visual speech information, particularly lip movements, offers a complementary modality for speech recognition.
- Developing robust lipreading systems is crucial for advancing human-computer interaction and assistive technologies.
Purpose of the Study:
- To design and implement a novel lipreading system capable of recognizing isolated words using solely visual information from lip movements.
- To evaluate the system's performance and accuracy in a controlled setting without acoustic data.
- To explore the efficacy of combining advanced image processing and machine learning techniques for visual speech recognition.
Main Methods:
- Utilized "snakes" for extracting visual features from the geometric space of lip movements.
- Applied Karhunen-Loeve transform (KLT) to identify principal components within the color eigenspace of lip images.
- Employed hidden Markov models (HMMs) for the sequential recognition of combined visual features.
Main Results:
- The lipreading system demonstrated high accuracy, achieving 94% recognition rate for ten isolated words.
- The system successfully performed word recognition using only color video of human lips, without any acoustic data.
- The combination of "snakes," KLT, and HMMs proved effective for visual feature extraction and sequence recognition.
Conclusions:
- Lipreading systems utilizing only visual information can achieve high accuracy in recognizing isolated words.
- The developed system offers a promising alternative or supplement to acoustic-based speech recognition, particularly in challenging acoustic environments.
- This research highlights the potential of advanced computer vision and machine learning techniques for robust visual speech recognition applications.
More Related Videos
10:11Portable Intermodal Preferential Looking (IPL): Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
Published on: December 14, 2012
06:07Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm
Published on: May 15, 2019