Related Experiment Video
Updated: Jun 3, 2026

10:25
Deep Learning-Based Segmentation of Cryo-Electron Tomograms
Published on: November 11, 2022
Learning diphone-based segmentation
Robert Daland1, Janet B Pierrehumbert
1Department of Linguistics, UCLA, Los Angeles, CA 90095-1543, USA. rdaland@humnet.ucla.edu
Cognitive Science
|March 25, 2011
Summary
This study develops a learnable model for word segmentation, showing infants can learn word boundaries from limited language exposure. The model
Area of Science:
- Computational Linguistics
- Developmental Psychology
- Speech Processing
Background:
- Previous diphone-based word segmentation models were considered unlearnable.
- Infant language acquisition relies on identifying word boundaries.
Purpose of the Study:
- To develop a statistically principled, learnable word segmentation model.
- To test the model's ability to recover phrase-medial word boundaries using real-world phonetic data.
- To investigate the impact of limited language exposure on segmentation performance.
Main Methods:
- Utilized Bayes' theorem and assumptions of infants' implicit knowledge.
- Developed unsupervised and semi-supervised learning models.
- Tested models on phonetic corpora from child-adult interactions.
Main Results:
- Achieved ceiling performance with 1 day to 1 month of language exposure.
- Demonstrated robustness to parameter and input representation variations.
- Observed undersegmentation in both learning and baseline models.
Conclusions:
- The developed model is learnable and effective for word segmentation.
- Limited language exposure is sufficient for infants to learn word boundaries.
- Undersegmentation has significant implications for overall speech processing.
