Related Experiment Video
Updated: Sep 4, 2025

Transmission of Multiple Signals through an Optical Fiber Using Wavefront Shaping
Published on: March 20, 2017
Modulation transfer functions for audiovisual speech
Nicolai F Pedersen1, Torsten Dau1, Lars Kai Hansen2
1Hearing Systems, Department of Health Technology, Technical University of Denmark, Kgs. Lyngby, Denmark.
Audiovisual speech synchrony reveals distinct facial motion patterns linked to speech rhythm. Natural speech shows correlations between acoustic envelope fluctuations and both mouth (3-4 Hz) and head (1-2 Hz) movements, suggesting a motor basis for speech timing.
Area of Science:
- Speech processing
- Auditory-visual perception
- Human-computer interaction
Background:
- Temporal synchrony between facial movements and speech acoustics is crucial for audiovisual speech perception.
- While low-frequency acoustic envelope fluctuations (below 10 Hz) correlate with facial motion, precise temporal relationships across different facial regions remain unclear.
Purpose of the Study:
- To investigate the precise temporal scales of synchronization between speech envelope modulations and facial motion in natural audiovisual speech.
- To identify distinct temporal patterns in facial movements corresponding to different speech rhythm scales.
Main Methods:
- Utilized regularized canonical correlation analysis (rCCA) to model speech envelope filters.
- Employed advanced video-based 3D facial landmark estimation on a large dataset (∼4000 speakers).
- Learned modulation transfer functions (MTFs) to correlate speech envelope with facial motion across speakers.
Main Results:
- Identified two distinct temporal scales of audiovisual speech synchrony: 3-4 Hz correlated with mouth movements and 1-2 Hz with global face/head motion.
- These timescales emerged specifically in natural audiovisual speech statistics across many speakers.
- Controlled speech tasks revealed only the 3-4 Hz modulations, suggesting slower rhythms are unique to natural speech.
Conclusions:
- Natural audiovisual speech exhibits distinct temporal regularities at syllable (3-4 Hz) and phrase (1-2 Hz) timescales, linked to specific facial motion patterns.
- The emergence of slower 1-2 Hz regularities only in crossmodal statistics suggests a potential motor origin for phrase-level speech timing.
More Related Videos
Related Concept Videos
Properties of Fourier Transform I
In radio broadcasting, multiple audio signals often need to be transmitted simultaneously. The Fourier...
State Space to Transfer Function
The transformation process begins with the state-space representation, characterized by the state equation and the output equation. These equations are typically represented as:
Network Function of a Circuit
Transfer Function to State Space
In an...
Transfer Function in Control Systems
To derive the transfer function, consider a general nth-order linear time-invariant...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...

