Related Experiment Video
Updated: Sep 13, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Robust frame-level speaker localization guided by multi-channel speech enhancement and inter-channel phase-difference
Shanmukha Srinivas Battula1, Hassan Taherian1, Ashutosh Pandey2
1Department of Computer Science and Engineering, The Ohio State University, Columbus, Ohio 43210, USA.
This study enhances speaker localization in noisy rooms using complex spectral mapping (CSM) for speech enhancement and direction-of-arrival (DOA) estimation. Multi-input multi-output (MIMO) systems with new loss functions significantly improve accuracy over single-output systems.
Area of Science:
- Signal Processing
- Acoustics
- Machine Learning
Background:
- Frame-level speaker localization is challenging in reverberant and noisy environments.
- Existing deep learning methods often process multi-channel audio directly, limiting performance.
- Accurate sound source localization relies on preserving inter-channel phase information.
Purpose of the Study:
- To develop a robust frame-level speaker localization method for adverse acoustic conditions.
- To investigate multi-channel speech enhancement using complex spectral mapping (CSM) for improved direction-of-arrival (DOA) estimation.
- To propose and evaluate novel multi-input multi-output (MIMO) systems and loss functions for enhanced localization accuracy.
Main Methods:
- Multi-channel speech enhancement using complex spectral mapping (CSM).
- Direction-of-arrival (DOA) estimation via weighted generalized cross-correlation with phase transform (GCC-PHAT).
- Comparison of multi-input single-output (MISO) and multi-input multi-output (MIMO) models, including new loss functions incorporating phase differences.
Main Results:
- CSM-derived phase estimates are reliable for frame-level DOA estimation.
- MIMO systems demonstrate superior performance compared to MISO systems.
- The proposed model achieves excellent frame-level speaker localization, outperforming related methods significantly.
Conclusions:
- The proposed CSM-based speech enhancement and DOA estimation framework effectively addresses limitations in noisy and reverberant environments.
- MIMO systems with novel phase-aware loss functions offer substantial improvements in speaker localization accuracy.
- The method achieves state-of-the-art frame-level localization performance, even surpassing utterance-level results of other approaches.
Related Concept Videos
Interference: Path Lengths
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Reconstruction of Signal using Interpolation
Distance Corrections
IR Spectrum Peak Splitting: Symmetric vs Asymmetric Vibrations

