Related Experiment Videos
Modeling the articulatory space using a hypercube codebook for acoustic-to-articulatory inversion
1Perceptual Science Laboratory, University of California Santa Cruz, Santa Cruz, California 95064, USA. slim@fuzzy.ucsc.edu
The Journal of the Acoustical Society of America
|August 27, 2005
Summary
This study introduces a novel acoustic-to-articulatory inversion method. It accurately reconstructs vocal tract shapes and dynamics, overcoming nonlinearities and nonuniqueness challenges in speech production.
Area of Science:
- Speech Science
- Acoustic Phonetics
- Signal Processing
Background:
- Acoustic-to-articulatory inversion is challenging due to nonlinearities and nonuniqueness between speech acoustics and vocal tract articulation.
- Existing methods often require excessive constraints, limiting the realistic recovery of vocal tract dynamics.
Purpose of the Study:
- To develop an inversion method that provides a complete description of possible solutions without over-constraining the system.
- To retrieve realistic temporal dynamics of vocal tract shapes from acoustic speech signals.
Main Methods:
- An adaptive sampling algorithm organizes a codebook into a hierarchy of hypercubes for near-constant acoustic resolution.
- Articulatory vectors are retrieved from the codebook, followed by nonlinear smoothing and regularization to recover the optimal articulatory trajectory.
Main Results:
- The developed inversion method precisely generates original formant trajectories from inverse articulatory parameters.
- It successfully recovers a realistic sequence of vocal tract shapes, demonstrating the method's effectiveness.
Conclusions:
- The proposed acoustic-to-articulatory inversion method effectively addresses the inherent complexities of speech production.
- It offers a robust approach for reconstructing vocal tract configurations and dynamics with high fidelity.