Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

293
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
293

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Rapid neutralization test based on ELISPOT for detecting antibodies against rhinovirus 1B.

Journal of microbiological methods·2026
Same author

The single-cell atlas of programmed cell death signature: A machine learning-based prognostic framework in breast cancer.

Journal of biomedical research·2026
Same author

A preliminary analysis of the inflammatory protein landscape in the CSF of mid- to late-stage Parkinson's disease: associations with motor severity and subtypes.

BMC neurology·2026
Same author

Polymer-Zn(II) sunscreens for protection against harmful blue ray.

Bioactive materials·2026
Same author

Prevalence and psychosocial interventions of mental disorders following traumatic spinal cord injury: a systematic review and meta-analysis based on knowledge graphs.

BMC psychiatry·2026
Same author

Tailoring the Porosity of Mesoporous Polyphenol Nanoparticles for Enhanced Photothermal Antibacterial Therapy.

Small (Weinheim an der Bergstrasse, Germany)·2026

Related Experiment Video

Updated: Aug 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K

Improving Speech Recognition Performance in Noisy Environments by Enhancing Lip Reading Accuracy.

Dengshi Li1, Yu Gao1, Chenyi Zhu1

  • 1School of Artificial Intelligence, Jianghan University, Wuhan 430056, China.

Sensors (Basel, Switzerland)
|February 28, 2023
PubMed
Summary

This study enhances noisy speech recognition by improving lip reading and cross-modal fusion. The novel approach significantly reduces word error rate (WER) in challenging acoustic conditions.

Keywords:
audiovisual speech recognitioncross-modal fusionlip readingnoisy environment

More Related Videos

Stimulating the Lip Motor Cortex with Transcranial Magnetic Stimulation
12:09

Stimulating the Lip Motor Cortex with Transcranial Magnetic Stimulation

Published on: June 14, 2014

19.1K
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

430

Related Experiment Videos

Last Updated: Aug 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K
Stimulating the Lip Motor Cortex with Transcranial Magnetic Stimulation
12:09

Stimulating the Lip Motor Cortex with Transcranial Magnetic Stimulation

Published on: June 14, 2014

19.1K
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

430

Area of Science:

  • Artificial Intelligence
  • Computer Science
  • Signal Processing

Background:

  • Speech recognition accuracy declines significantly in noisy environments.
  • Lip reading (visual information) offers a noise-invariant alternative to improve speech recognition.
  • Cross-modal fusion of audio and visual data is crucial for robust speech recognition.

Purpose of the Study:

  • To enhance speech recognition accuracy in noisy environments.
  • To improve lip reading performance and the effectiveness of cross-modal fusion.
  • To develop a model that leverages lip movements for better audio interpretation.

Main Methods:

  • Constructed a one-to-many lip-to-speech mapping model to interpret lip movements.
  • Preserved audio representations by modeling inter-relationships between audiovisual data.
  • Employed a joint cross-fusion model with attention mechanisms for intermodal relationship exploitation.
  • Calculated cross-attention weights based on feature correlations between modalities.

Main Results:

  • Achieved a 4.0% reduction in word error rate (WER) in a -15 dB SNR environment compared to the baseline.
  • Demonstrated a 10.1% reduction in WER compared to traditional speech recognition.
  • Showcased significant performance improvements in various noisy conditions.

Conclusions:

  • The proposed method effectively improves speech recognition in noisy environments.
  • Integrating enhanced lip reading with cross-modal fusion offers a robust solution.
  • The model's ability to extract audio from video input is a key advancement.