Related Experiment Videos
Inferring articulation and recognizing gestures from acoustics with a neural network trained on x-ray microbeam data
G Papcun1, J Hochberg, T R Thomas
1Computing and Communications Division, Los Alamos National Laboratory, New Mexico 87545.
The Journal of the Acoustical Society of America
|August 1, 1992
Summary
This study introduces a novel neural network method to infer speech articulatory parameters from acoustic data. The model accurately predicts vocal tract movements, advancing speech research and technology.
Area of Science:
- Speech Science
- Computational Linguistics
- Machine Learning
Background:
- Articulatory speech synthesis requires accurate vocal tract movement data.
- Inferring articulatory parameters directly from acoustics is challenging.
Purpose of the Study:
- To develop and evaluate a neural network for inferring articulatory parameters from acoustic speech signals.
- To assess the model's performance in both learning and generalization conditions.
Main Methods:
- Trained a neural network on paired x-ray microbeam (articulatory) and acoustic data from three speakers.
- Used vertical movements of the lower lip, tongue tip, and tongue dorsum for English stop consonants.
- Evaluated model accuracy using root-mean-square error, correlation, and gesture recognition rates.
Main Results:
- Inferred articulatory trajectories closely matched actual movements (learning and generalization).
- Achieved high gesture recognition rates (94.4%-98.9%) in learning and cross-speaker conditions.
- Demonstrated 75% accuracy for consonants outside the training set, showing good generalization.
Conclusions:
- The neural network effectively infers articulatory parameters from acoustics.
- The method shows strong potential for applications in speech synthesis and analysis.
- Articulator movement regularity is linked to consonant formation criticality.