Related Experiment Video
Updated: Oct 2, 2025

08:45
Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
14.8K
A Speech-Level-Based Segmented Model to Decode the Dynamic Auditory Attention States in the Competing Speaker Scenes
Lei Wang1,2, Yihan Wang1, Zhixing Liu1
1Department of Electrical and Electronic Engineering, Southern University of Science and Technology, Shenzhen, China.
Frontiers in Neuroscience
|February 28, 2022
Summary
A new speech-RMS-level-based model significantly improves auditory attention decoding (AAD) in noisy environments. This segmented model enhances performance for both sustained and switched attention, outperforming unified models for reliable speech perception.
Area of Science:
- Neuroscience
- Auditory Perception
- Signal Processing
Background:
- Human listeners in competing speaker environments must dynamically manage auditory attention.
- Cortical tracking of speech envelope is key for decoding target speech from neural signals.
- Previous research indicated root mean square (RMS)-level-based speech segmentation aids target speech perception and sustained attention.
Purpose of the Study:
- To investigate the impact of RMS-level-based speech segmentation on auditory attention decoding (AAD) performance.
- To evaluate AAD with sustained and switched attention in competing speaker scenarios.
- To develop objective biomarkers from cortical activity for dynamic auditory attention states.
Main Methods:
- Subjects focused on or switched attention between two competing speech streams.
- Neural responses to higher- and lower-RMS-level speech segments were analyzed using linear temporal response functions (TRF).
- AAD performance of a unified TRF model was compared against an RMS-level-based segmented decoding model.
Main Results:
- TRF component weights at approximately 100-ms lag were sensitive to auditory attention switching.
- The segmented AAD model demonstrated improved attention decoding performance over the unified model under sustained and switched attention.
- The segmented model showed robust decoding across a wide range of signal-to-masker ratios (SMRs) and short decision windows.
Conclusions:
- TRF weights and AAD accuracy serve as effective indicators for detecting dynamic changes in auditory attention.
- RMS-level-based speech segmentation enhances auditory attention decoding in complex auditory scenes.
- The segmented AAD model shows potential for decoding dynamic attention states in realistic auditory environments.

