Related Experiment Video
Updated: May 10, 2026

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
ASEAF: attention-SincNet driven EEG-audio fused target speaker extraction network
Yuhang Yang1, Yuan Liao2, Qiushi Han1
1College of electronic and optical engineering & college of flexible electronics (future technology), Nanjing University of Posts and Telecommunications, Jiangsu 210023, People's Republic of China.
Abstract:
This study addresses the challenge of selective auditory attention in noisy environments by proposing an electroencephalography (EEG)-based target speaker extraction model, ASEAF, designed to mimic neural decoding through tailored spatio-temporal feature extraction and cross-modal fusion. The model achieves precise extraction of the target speaker's speech by simultaneously processing EEG and audio signals. ASEAF comprises four modules: an EEG encoder using CNN and self-attention for spatio-temporal features, an audio encoder with SincNet for frequency-aware processing, a dual-path LSTM speaker extractor for fused feature masking, and a CNN decoder for waveform reconstruction. This innovative integration advances neural-signal-based speech reconstruction by providing insights into cross-modal interactions. Experiments on the Cocktail Party dataset, KUL dataset and DTU dataset demonstrate that ASEAF outperforms state-of-the-art models across multiple metrics, with an average improvement of 11.5% in scale-invariant signal-to-distortion ratio improvement. This work offers a more effective hearing aid solution for individuals with hearing impairments and advances the field of brain-computer interfaces.

