Related Experiment Video
Updated: Jun 16, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Unsupervised speech segmentation: an analysis of the hypothesized phone boundaries
Odette Scharenborg1, Vincent Wan, Mirjam Ernestus
1Centre for Language and Speech Technology, Radboud University Nijmegen, Erasmusplein 1, 6525 HT Nijmegen, The Netherlands. o.scharenborg@let.ru.nl
The Journal of the Acoustical Society of America
|February 9, 2010
Summary
Unsupervised speech segmentation methods show promise, but struggle to match human transcriptions. Improving these methods requires combining bottom-up acoustic analysis with top-down information for better phone boundary detection.
Area of Science:
- Speech Processing
- Computational Linguistics
- Machine Learning
Background:
- Unsupervised automatic phone segmentation methods often yield similar performance in boundary detection.
- These methods do not perfectly replicate human-created reference transcriptions.
Purpose of the Study:
- Investigate fundamental issues in unsupervised segmentation algorithms.
- Compare acoustic-only segmentation with human reference transcriptions.
Main Methods:
- Analyzed an unsupervised speech segmentation method using acoustic change detection.
- Compared hypothesized boundaries with human-annotated segment boundaries.
- Performed statistical analyses on segmentation errors.
Main Results:
- Acoustic change is a reliable indicator for segment boundaries, with over two-thirds accuracy.
- Errors correlate with segment duration, similar segment sequences, and dynamic phones.
- Current one-stage methods need enhancement into two-stage approaches.
Conclusions:
- Enhancing unsupervised segmentation requires integrating bottom-up and top-down information.
- Two-stage methods can improve accuracy while maintaining flexibility and language independence.

