Related Experiment Video
Updated: May 16, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Automatic measurement of voice onset time using discriminative structured prediction
Morgan Sonderegger1, Joseph Keshet
1Department of Computer Science and Department of Linguistics, University of Chicago, 1100 E 58th Street, Chicago, Illinois 60637, USA. morgan.sonderegger@mcgill.ca
This study introduces an automatic algorithm for measuring voice onset time (VOT) in speech. The method achieves human-level accuracy and is robust across different speech corpora and with limited training data.
Area of Science:
- Speech processing
- Acoustic phonetics
- Machine learning
Background:
- Voice onset time (VOT) is a crucial phonetic feature for distinguishing speech sounds.
- Manual measurement of VOT is time-consuming and subject to inter-annotator variability.
- Automatic methods for VOT measurement are needed for large-scale speech analysis.
Purpose of the Study:
- To develop a discriminative large-margin algorithm for automatic VOT measurement.
- To predict VOT as a structured output from speech segments.
- To achieve performance comparable to human annotators.
Main Methods:
- A large-margin algorithm was trained using manually labeled speech data.
- The algorithm predicts VOT from acoustic features derived from spectral and temporal cues.
- Feature functions mimic those used by human VOT annotators.
Main Results:
- The algorithm achieved performance near human inter-transcriber reliability across four speech corpora.
- The method demonstrated favorable comparison with previous automatic VOT measurement techniques.
- Performance remained consistent when training and testing on different corpora.
Conclusions:
- The developed algorithm offers a practical and accurate solution for automatic VOT measurement.
- The method shows robustness to variations in speech corpora and requires minimal training data (50-250 examples).
- This approach has significant potential for applications in speech research and technology.
More Related Videos
05:48Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
06:09P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023