Related Experiment Videos
Use of syllable-scale timing to discriminate words
R F Port1, W T Reilly, D P Maki
1Indiana University, Bloomington 47405.
The Journal of the Acoustical Society of America
|January 1, 1988
Summary
This study shows that acoustic boundary timing in speech can identify words, even with imperfect phonetic information. This temporal information significantly improves word recognition accuracy beyond chance levels.
Area of Science:
- Speech processing
- Acoustic phonetics
- Machine learning for speech recognition
Background:
- Automatic speech recognition often relies on detailed phonetic information.
- Primitive phonetic labeling presents a challenge for accurate word identification.
- Understanding the role of temporal acoustic cues is crucial for robust speech processing.
Purpose of the Study:
- To investigate the effectiveness of acoustic boundary timing for word recognition with imperfect phonetic labeling.
- To determine if temporal information can improve word identification accuracy.
- To assess the impact of phonetic confusions on word recognition performance.
Main Methods:
- Utilized discriminant analysis to combine acoustic boundary locations.
- Selected sets of two-syllable words differing in stress, vowel tensify, and segmental identity.
- Analyzed speech data from multiple talkers, speaking tempos, and using sound spectrograms.
- Tested recognition accuracy on unseen speech productions.
Main Results:
- Achieved word recognition accuracy 6.3 times better than chance using six segmental boundaries.
- Maintained performance five times better than chance even with more confusable words.
- Demonstrated reasonable word discrimination using a subset of temporal variables.
Conclusions:
- Segmental timing provides significant information for word identification, even with weak phonetic labeling.
- Temporal acoustic cues are valuable for enhancing speech recognition systems.
- This approach offers a promising direction for developing more robust automatic speech recognition.