Related Experiment Video
Updated: Oct 10, 2025

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
An Effective Conversion of Visemes to Words for High-Performance Automatic Lipreading
Souheil Fenghour1, Daqing Chen1, Kun Guo2
1School of Engineering, London South Bank University, London SE1 0AA, UK.
Viseme-to-word conversion in lipreading is challenging due to homophemes. A novel deep learning model significantly improves sentence prediction accuracy by effectively distinguishing between similar-sounding words.
Area of Science:
- Computer Science
- Artificial Intelligence
- Signal Processing
Background:
- Viseme-based lipreading systems show promise for sentence decoding.
- System performance is bottlenecked by inefficient viseme-to-word conversion.
- Homophenes, where visemes map to multiple words, cause significant accuracy drops.
Purpose of the Study:
- To develop an efficient deep learning model for viseme-to-word conversion.
- To address the challenge of homophenes in lipreading.
- To improve the accuracy of spoken sentence prediction from videos.
Main Methods:
- A deep learning network model utilizing an Attention-based Gated Recurrent Unit (AGRU) was proposed.
- The AGRU model was compared against three alternative approaches.
- Experiments were conducted on the LRS2 and LRS3 benchmark datasets.
Main Results:
- The proposed AGRU model demonstrated robustness, high efficiency, and short execution times.
- The model effectively converted visemes to words and discriminated between homophenes.
- Achieved a word accuracy rate of 79.6% on the LRS2 dataset, a 15.0% improvement over state-of-the-art.
Conclusions:
- The developed AGRU model significantly enhances viseme-to-word conversion accuracy in lipreading.
- The model's efficiency and robustness make it suitable for practical applications.
- This approach offers a substantial improvement for predicting spoken sentences from visual data.
More Related Videos
05:38Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Improving Translational Accuracy
Components of Language
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Genetic Lingo
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...