Related Experiment Video
Updated: Oct 15, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Spatial alignment between faces and voices improves selective attention to audio-visual speech
Justin T Fleming1, Ross K Maddox2, Barbara G Shinn-Cunningham3
1Speech and Hearing Bioscience and Technology Program, Harvard University, 243 Charles Street, Boston, Massachusetts 02114, USA.
Spatial alignment of faces and voices enhances speech understanding in noisy, multi-talker settings. This visual-auditory (AV) spatial alignment is crucial for effectively focusing on a target speaker amidst distractions.
Area of Science:
- Auditory Perception
- Visual Perception
- Human-Computer Interaction
Background:
- Visual cues, such as seeing a talker's face, significantly improve speech intelligibility in noisy environments when audio-visual signals are temporally synchronized.
- The role of spatial alignment between faces and voices in multi-talker scenarios remains less understood.
Purpose of the Study:
- To investigate the impact of spatial alignment between faces and voices on selective attention to speech in noise.
- To determine if spatial alignment is a critical factor for audio-visual (AV) speech perception in complex listening environments.
Main Methods:
- Conducted online experiments using a selective attention task in multi-talker noise.
- Manipulated spatial alignment of faces and voices (same vs. different hemifield) and signal-to-noise ratio (SNR).
- Investigated the influence of eye gaze direction on performance in auditory-only conditions.
Main Results:
- Task performance improved when faces and voices were spatially aligned (same hemifield).
- Spatial misalignment between audio and visual speech signals incurred a performance cost.
- The effect of AV spatial alignment was more pronounced at lower SNRs, though lipreading accuracy introduced a floor effect.
Conclusions:
- Spatial alignment between faces and voices is a significant factor contributing to the ability to selectively attend to audio-visual (AV) speech.
- Effective AV speech perception in noisy, multi-talker environments relies on both temporal and spatial congruence of speech cues.
More Related Videos
08:45Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
11:15fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
Related Concept Videos
Association Areas of the Cortex
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Auditory Perception
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Facial Feedback Hypothesis
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
The Anchoring-and-Adjustment Heuristic