Related Experiment Video
Updated: Nov 7, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
A dual-stream deep attractor network with multi-domain learning for speech dereverberation and separation
Hangting Chen1, Pengyuan Zhang1
1Key Laboratory of Speech Acoustics & Content Understanding, Institute of Acoustics, Chinese Academy of Sciences, China; University of Chinese Academy of Sciences, Beijing, China.
This study introduces a dual-stream deep attractor network (DAN) for improved speech separation and dereverberation in noisy environments. The novel approach enhances signal quality in reverberant conditions, outperforming existing methods.
Area of Science:
- Speech processing
- Machine learning
- Audio signal analysis
Background:
- Deep attractor networks (DANs) excel at speech separation using discriminative embeddings but struggle in reverberant environments.
- Existing methods like permutation invariant training (PIT) have limitations in complex acoustic conditions.
Purpose of the Study:
- To develop an improved speech separation and dereverberation method for reverberant environments.
- To enhance the performance of deep attractor networks (DANs) in challenging acoustic conditions with multiple speakers.
Main Methods:
- A novel dual-stream DAN with multi-domain learning was proposed, integrating a speaker encoding stream (SES) and a speech decoding stream (SDS).
- The SES models speaker information in a Fourier transform-based embedding space, while the SDS estimates the early sound component in the time domain.
- Clustering losses were employed to refine attractor estimation.
Main Results:
- The dual-stream DAN achieved significant scale-invariant source-to-distortion ratio (SI-SDR) improvements: 9.8 dB for 2-speaker and 7.5 dB for 3-speaker evaluations.
- Performance gains exceeded the baseline DAN by 2.0/0.7 dB and convolutional time-domain audio separation network (Conv-TasNet) by 1.0/0.5 dB on reverberant datasets.
- The method demonstrated effectiveness in handling variable numbers of speakers in reverberant settings.
Conclusions:
- The proposed dual-stream DAN effectively addresses speech separation and dereverberation challenges in reverberant environments.
- This multi-domain learning approach offers a substantial improvement over existing speech separation techniques.
- The method shows promise for real-world applications requiring robust speech processing in adverse acoustic conditions.
Related Concept Videos
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Air-entraining Agents
Uniform Depth Channel Flow: Problem Solving
Multi-input and Multi-variable systems
In the absence of...

