Related Experiment Video
Updated: Aug 15, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improved Speech Spatial Covariance Matrix Estimation for Online Multi-Microphone Speech Enhancement.
Minseung Kim1, Sein Cheong1, Hyungchan Song1
1School of Electrical Engineering and Computer Science, Gwangju Institute of Science and Technology, Buk-gu, Gwangju 61005, Republic of Korea.
This study introduces an improved method for online multi-microphone speech enhancement, focusing on better estimation of speech spatial covariance matrices (SCMs). The new approach enhances speech clarity and intelligibility in noisy environments.
Area of Science:
- Signal Processing
- Acoustics
- Machine Learning
Background:
- Online multi-microphone speech enhancement requires causal estimation of acoustic parameters.
- Accurate estimation of speech spatial covariance matrices (SCMs) is crucial for performance.
Purpose of the Study:
- To propose an improved causal estimator for speech SCMs in online multi-microphone speech enhancement.
- To enhance the accuracy of estimating speech power spectral density (PSD) and relative transfer function (RTF).
Main Methods:
- Utilized temporal cepstrum smoothing (TCS) for speech PSD estimation.
- Developed a novel RTF estimator based on time difference of arrival (TDoA) from cross-correlation.
- Refined speech SCM estimates using clean speech spectrum and power spectrum information.
Main Results:
- The proposed method demonstrated superior performance on the CHiME-4 database.
- Improvements were measured using perceptual evaluation of speech quality (PESQ), extended short-time objective intelligibility (eSTOI), and scale-invariant signal-to-distortion ratio (SISDR).
Conclusions:
- The novel approach significantly enhances online speech enhancement performance.
- The refined SCM estimation leads to better speech quality and intelligibility in noisy conditions.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
08:45Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Reconstruction of Signal using Interpolation
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Unsoundness of Aggregate due to Volume Change