Related Experiment Video
Updated: May 27, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
A subspace approach based on embedded prewhitening for voice activity detection
Dong Kook Kim1, Joon-Hyuk Chang
1School of Electronic and Computer Engineering, Chonnam National University, Yongbong, 500-757 Gwangju, Republic of Korea. dkim@chonnam.ac.kr
This study introduces a novel subspace approach for voice activity detection (VAD). This method enhances performance, especially in noisy environments, outperforming traditional Gaussian models.
Area of Science:
- Signal Processing
- Speech Recognition
- Acoustics
Background:
- Voice Activity Detection (VAD) is crucial for speech processing systems.
- Traditional VAD methods can struggle in low signal-to-noise ratio (SNR) conditions.
- Subspace methods offer potential for improved VAD performance.
Purpose of the Study:
- To propose a new subspace approach for Voice Activity Detection (VAD).
- To evaluate the performance of the proposed VAD algorithm against existing methods.
- To demonstrate the effectiveness of the approach in challenging acoustic environments.
Main Methods:
- A subspace approach utilizing an embedded prewhitening scheme.
- Simultaneous diagonalization of clean speech and noise covariance matrices.
- A decision rule based on the likelihood ratio test in the signal subspace domain.
Main Results:
- The proposed subspace-based VAD algorithm demonstrates superior performance.
- Outperforms conventional Gaussian model-based methods in low SNR conditions.
- Effective in distinguishing speech from noise in challenging environments.
Conclusions:
- The subspace approach offers a robust solution for VAD.
- The embedded prewhitening scheme enhances VAD accuracy.
- This method is particularly beneficial for applications with low signal quality.
Related Concept Videos
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same frequency...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...

