Related Experiment Video
Updated: May 16, 2026

12:43
A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
Published on: February 21, 2011
A procedure for estimating gestural scores from speech acoustics.
Hosung Nam1, Vikramjit Mitra, Mark Tiede
1Haskins Laboratories, 300 George Street, Suite 900, New Haven, Connecticut 06511, USA. nam@haskins.yale.edu
The Journal of the Acoustical Society of America
|December 13, 2012
Summary
This study introduces a new method for annotating speech gestures using an iterative time-warping technique. This approach reliably maps vocal tract actions in speech, improving speech dataset analysis.
Area of Science:
- Phonetics and Speech Science
- Computational Linguistics
- Speech Technology
Background:
- Speech is characterized by vocal tract actions (gestures) organized in temporal patterns (gestural scores).
- Existing speech datasets lack formal gestural annotations, hindering detailed analysis.
- A standardized procedure for gestural annotation is currently unavailable.
Purpose of the Study:
- To develop and validate an automated procedure for gestural annotation of natural speech.
- To create a method for generating temporally optimized gestural scores for speech utterances.
- To enhance the usability of speech datasets through reliable gestural information.
Main Methods:
- An iterative analysis-by-synthesis architecture was designed for gestural annotation.
- The Haskins Laboratories Task Dynamics and Application (TADA) model generated prototype gestural scores.
- Time-warping optimized gestural scores by minimizing acoustic distance between original and synthesized speech.
Main Results:
- The proposed iterative time-warping approach achieved superior performance compared to conventional methods.
- Reliable gestural annotations were successfully generated for natural speech datasets.
- The method provides a robust way to link acoustic speech signals with articulatory gestures.
Conclusions:
- The developed iterative gestural annotation architecture offers a reliable and effective solution.
- This method addresses the current gap in gestural annotation for speech datasets.
- The findings facilitate more in-depth analysis of speech production and acoustics.

