Related Experiment Video
Updated: Sep 13, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Robust frame-level speaker localization guided by multi-channel speech enhancement and inter-channel phase-difference
Shanmukha Srinivas Battula1, Hassan Taherian1, Ashutosh Pandey2
1Department of Computer Science and Engineering, The Ohio State University, Columbus, Ohio 43210, USA.
Abstract:
In the presence of room reverberation and background noise, the performance of frame-level speaker localization is severely limited. To address this challenge, this study performs multi-channel speech enhancement based on complex spectral mapping (CSM), followed by direction-of-arrival (DOA) estimation using weighted generalized cross-correlation with phase transform (GCC-PHAT). The proposed approach differs from prevailing deep learning methods that operate on multi-channel inputs directly for speaker localization. This study initially investigates multi-input single-output (MISO) based speech enhancement models with two loss functions and then extends to multi-input multi-output (MIMO) based modeling for conceptual and computational efficiency. The results demonstrate that the phase estimates obtained from CSM models are reliable for frame-level DOA estimation and MIMO systems outperform MISO systems. In addition, the study proposes new multi-channel loss functions for MIMO systems that incorporate phase differences in order to better preserve inter-channel phase relations, which is key to accurate sound localization. Systematic evaluations with multiple microphone array geometries using both simulated and recorded room impulse responses, as well as real recordings, demonstrate that the proposed model yields excellent frame-level speaker localization results in reverberant and noisy environments and outperforms related methods by a large margin, even surpassing their utterance-level results.
Related Concept Videos
Interference: Path Lengths
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Reconstruction of Signal using Interpolation
Distance Corrections
IR Spectrum Peak Splitting: Symmetric vs Asymmetric Vibrations

