Related Experiment Video
Updated: Jun 7, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Leveraging laryngograph data for robust voicing detection in speech
Yixuan Zhang1, Heming Wang1, DeLiang Wang1,2
1Department of Computer Science and Engineering, The Ohio State University, Columbus, Ohio 43210, USA.
This study introduces a novel supervised voicing detection model using laryngograph data. The advanced CrossNet-based model achieves robust speech signal analysis, outperforming existing methods and generalizing to new datasets.
Area of Science:
- Speech processing
- Machine learning
- Bioacoustics
Background:
- Accurate voiced interval detection is crucial for speech analysis, including pitch tracking.
- Existing methods often require dataset-specific parameter tuning and exhibit limited generalization.
- Conventional signal processing and deep learning approaches face challenges in real-world speech applications.
Purpose of the Study:
- To develop a robust supervised voicing detection model for speech signals.
- To improve the generalization capability of voicing detection models.
- To provide a reliable alternative to conventional methods with less parameter tuning.
Main Methods:
- A supervised voicing detection model adapted from the CrossNet architecture was developed.
- The model was trained using reference voicing decisions from laryngograph datasets.
- Pretraining strategies were investigated to enhance model generalization.
Main Results:
- The proposed model demonstrated robust voicing detection performance.
- It significantly outperformed established baseline methods in accuracy.
- The model showed excellent generalization to unseen speech datasets.
Conclusions:
- The developed supervised voicing detection model offers superior performance and generalization.
- Leveraging laryngograph data and CrossNet architecture addresses limitations of prior methods.
- The provided source code and datasets will aid future research in speech signal processing.
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Assessment of Ventilation II: Respiratory Depth and Rhythm
Respiratory depth measures the volume of air inhaled or exhaled during a breath. It can vary from shallow to deep and typically remains consistent when a person is at rest or asleep. Occasionally, individuals will automatically inhale deeply, known as sighing, which inflates the lungs with more air than normal breathing.
To assess respiratory depth, observe the degree of chest excursion or movement:
Physical Assessment of the Respiratory Tract IV: Auscultation
Breath Sounds
Breath sounds are categorized into vesicular, bronchovesicular, and bronchial.
Suctioning the Oropharyngeal Airway
After assembling the equipment, the nurse should practice hand hygiene and don appropriate PPE according to infection control guidelines to avoid the...
Respiratory System Abnormal Finding II: Palpation and Auscultation
Palpation Findings
During a respiratory assessment, palpation can reveal several vital abnormalities:

