Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Neural Circuits01:25

Neural Circuits

1.8K
Neural circuits and neuronal pools are two of the main structures found in the nervous system. Neural circuits are networks of neurons that work together to carry out a specific task or process. They consist of interconnected neurons and glial cells, which provide structural and metabolic support.
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
1.8K
Parallel Processing01:20

Parallel Processing

334
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
334

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Hybrid Deep Neural Network Framework Combining Skeleton and Gait Features for Pathological Gait Recognition.

Bioengineering (Basel, Switzerland)·2023
Same author

A Deep Learning-Based Semantic Segmentation Model Using MCNN and Attention Layer for Human Activity Recognition.

Sensors (Basel, Switzerland)·2023
Same author

Deep-Learning-Based ADHD Classification Using Children's Skeleton Data Acquired through the ADHD Screening Game.

Sensors (Basel, Switzerland)·2023
Same author

Deep Learning-Based ADHD and ADHD-RISK Classification Technology through the Recognition of Children's Abnormal Behaviors during the Robot-Led ADHD Screening Game.

Sensors (Basel, Switzerland)·2023
Same author

A Low-Cost Foot-Placed UWB and IMU Fusion-Based Indoor Pedestrian Tracking System for IoT Applications.

Sensors (Basel, Switzerland)·2022
Same author

Noise-Robust Multimodal Audio-Visual Speech Recognition System for Speech-Based Interaction Applications.

Sensors (Basel, Switzerland)·2022

Related Experiment Video

Updated: Oct 7, 2025

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
05:38

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology

Published on: June 29, 2021

2.5K

Lipreading Architecture Based on Multiple Convolutional Neural Networks for Sentence-Level Visual Speech Recognition.

Sanghun Jeon1, Ahmed Elsharkawy1, Mun Sang Kim1

  • 1Center for Healthcare Robotics, Gwangju Institute of Science and Technology (GIST), School of Integrated Technology, Gwangju 61005, Korea.

Sensors (Basel, Switzerland)
|January 11, 2022
PubMed
Summary

This study introduces a new deep learning model for visual speech recognition (VSR) that improves accuracy in transcribing speech from lip movements. The advanced architecture effectively handles challenges like homophones and short words, enhancing VSR reliability.

Keywords:
3D densely connected CNN3D multi-layer feature fusion CNNconvolutional neural networkdeep learninglipreadingspeech recognitionvisual speech recognition

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

674

Related Experiment Videos

Last Updated: Oct 7, 2025

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
05:38

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology

Published on: June 29, 2021

2.5K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

674

Area of Science:

  • Computer Science
  • Artificial Intelligence
  • Machine Learning

Background:

  • Visual Speech Recognition (VSR) uses lip movements for speech transcription, with deep learning showing high accuracy.
  • Existing VSR systems struggle with homophones (similar-sounding words) and short words due to visual ambiguity and insufficient data.
  • Limitations in traditional VSR hinder practical applications, necessitating improved models.

Purpose of the Study:

  • To propose a novel lipreading architecture designed to overcome the limitations of current VSR systems.
  • To enhance the accuracy and reliability of visual speech recognition, particularly in challenging scenarios.
  • To reduce character and word error rates in VSR tasks.

Main Methods:

  • A new lipreading architecture combining three types of 3D Convolutional Neural Networks (CNNs): 3D CNN, densely connected 3D CNN, and multi-layer feature fusion 3D CNN.
  • The CNNs are followed by a two-layer bi-directional gated recurrent unit.
  • The entire network was trained using connectionist temporal classification.

Main Results:

  • The proposed architecture significantly reduced character error rates by 5.681% and word error rates by 11.282% on an unseen-speaker dataset.
  • The model demonstrated improved performance in distinguishing homophones and recognizing short words.
  • The VSR system showed increased reliability in practical applications, even with visual ambiguity.

Conclusions:

  • The novel deep learning architecture effectively addresses key challenges in visual speech recognition, such as word ambiguity and data limitations.
  • The proposed method offers a significant improvement over baseline models, enhancing the accuracy and robustness of VSR.
  • This research contributes to more reliable and practical visual speech recognition systems.