Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Classification of Signals01:30

Classification of Signals

In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Chunking and Rehearsal in Sensory Memory01:22

Chunking and Rehearsal in Sensory Memory

Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of information more...
Automatic Processing and Automatic Social Behavior01:28

Automatic Processing and Automatic Social Behavior

Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
Auditory Perception01:17

Auditory Perception

The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the cochlea, a...
Force Classification01:22

Force Classification

Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multi-channel auto-encoders for learning domain invariant representations enabling superior classification of histopathology images.

Medical image analysis·2022
Same author

Matching Larger Image Areas for Unconstrained Face Identification.

IEEE transactions on cybernetics·2018
Same author

Largest Matching Areas for Illumination and Occlusion Robust Face Recognition.

IEEE transactions on cybernetics·2016
Same author

Target therapy of multiple myeloma by PTX-NPs and ABCG2 antibody in a mouse xenograft model.

Oncotarget·2015

Related Experiment Video

Updated: May 10, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Robust audio-visual speech recognition under noisy audio-video conditions.

Darryl Stewart, Rowan Seymour, Adrian Pass

    IEEE Transactions on Cybernetics
    |June 13, 2013
    PubMed
    Summary

    The Maximum Weighted Stream Posterior (MWSP) model offers robust audio-visual speech recognition by dynamically weighting audio and video streams, even with unknown noise. This modality-independent method achieves excellent performance across various corruption types and levels.

    More Related Videos

    An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
    07:52

    An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

    Published on: March 13, 2026

    Related Experiment Videos

    Last Updated: May 10, 2026

    Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
    05:48

    Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

    Published on: August 9, 2024

    An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
    07:52

    An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

    Published on: March 13, 2026

    Area of Science:

    • Speech Recognition
    • Signal Processing
    • Machine Learning

    Background:

    • Audio-visual speech recognition systems often struggle with noisy or corrupted audio/video streams.
    • Existing methods may require specific signal measurements or prior knowledge of noise characteristics.

    Purpose of the Study:

    • To introduce and evaluate the Maximum Weighted Stream Posterior (MWSP) model for robust audio-visual speech recognition.
    • To demonstrate the model's effectiveness in environments with unknown and time-varying stream corruption.
    • To highlight the modality-independent nature of MWSP.

    Main Methods:

    • Developed the Maximum Weighted Stream Posterior (MWSP) model for dynamic stream weighting.
    • Evaluated MWSP on the XM2VTS database for speaker-independent audio-visual speech recognition.
    • Introduced various types and levels of corruption to audio and video streams.

    Main Results:

    • MWSP demonstrated excellent recognition performance compared to dynamic and fixed-weighted approaches.
    • The model dynamically adjusted stream weights frame-by-frame based on noise levels and modality reliability.
    • Robust performance was maintained across all tested conditions, including clean and corrupted streams.

    Conclusions:

    • MWSP provides a robust and efficient solution for audio-visual speech recognition under challenging environmental conditions.
    • The model's modality-independent nature and dynamic weighting mechanism eliminate the need for prior noise knowledge.
    • MWSP offers a significant advancement in handling corrupted audio and video streams for speech recognition.