Related Experiment Video
Updated: Jan 15, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
ctPuLSE: Close-talk, and pseudo-label based far-field, speech enhancement
1Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen 518055, Guangdong, People's Republic of China.
This study introduces ctPuLSE, a novel method for training neural speech enhancement models directly on real-world recordings. It uses close-talk speech enhancement to create pseudo-labels for improving far-field speech enhancement generalizability.
Area of Science:
- Speech Processing
- Artificial Intelligence
- Machine Learning
Background:
- Current neural speech enhancement models rely on simulated data, limiting real-world performance.
- Training directly on real mixtures is challenging due to the lack of clean speech supervision.
Purpose of the Study:
- To develop a method for training speech enhancement models directly on real-recorded mixtures.
- To improve the generalizability of far-field speech enhancement models to real-world conditions.
Main Methods:
- Propose ctPuLSE: train an enhancement model on simulated data to process real close-talk mixtures.
- Utilize the enhanced close-talk speech as pseudo-labels for training far-field enhancement models on real paired mixtures.
Main Results:
- ctPuLSE effectively generates high-quality pseudo-labels from real close-talk mixtures.
- The proposed method significantly enhances the generalizability of far-field speech enhancement models on real data.
Conclusions:
- Close-talk pseudo-labeling offers a viable solution for supervised training on real-world speech data.
- ctPuLSE demonstrates a promising approach for robust neural speech enhancement in practical scenarios.
Related Concept Videos
Echo
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Non-Verbal Cues
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
Interference: Path Lengths
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....

