Related Experiment Video
Updated: Apr 27, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
A hierarchical framework approach for voice activity detection and speech enhancement
Yan Zhang1, Zhen-min Tang2, Yan-ping Li3
1College of Computer Science and Technology, Nanjing University of Science and Technology (NUST), Nanjing 210094, China ; College of Information Technology, Jinling Institute of Technology (JIT), Nanjing 211169, China.
This study introduces a novel hierarchical framework for voice activity detection (VAD) and speech enhancement. The proposed method effectively reduces noise and improves accuracy in challenging noisy conditions.
Area of Science:
- Signal Processing
- Speech Technology
- Machine Learning
Background:
- Voice Activity Detection (VAD) is crucial for speech and speaker recognition systems.
- Existing VAD techniques struggle with performance in noisy environments.
- Speech enhancement is often required to improve the quality of degraded speech signals.
Purpose of the Study:
- To propose a hierarchical framework integrating VAD and speech enhancement.
- To improve the accuracy and robustness of VAD in noisy conditions.
- To evaluate the effectiveness of the proposed approach against established methods.
Main Methods:
- A hierarchical framework combining VAD and speech enhancement.
- Modified Wiener Filter (MWF) for noise reduction in speech enhancement.
- Feature selection and voting mechanism for reliable VAD decision-making.
Main Results:
- The proposed hierarchical framework demonstrates strong performance.
- The method is effective across various noisy conditions.
- Experimental results validate the approach using TIMIT and NOISEX-92 databases.
Conclusions:
- The integrated VAD and speech enhancement framework offers a robust solution.
- The proposed method shows significant improvements in noisy environments.
- This approach advances the field of robust speech processing.
More Related Videos
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Design Example
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Amplifying Signals via Enzymatic Cascade

