Related Experiment Video
Updated: Dec 31, 2025

10:16
Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
Published on: December 2, 2011
14.4K
A modular architecture for articulatory synthesis from gestural specification.
Rachel Alexander1, Tanner Sorensen1, Asterios Toutios1
1Signal Analysis & Interpretation Laboratory (SAIL), University of Southern California, Los Angeles, California 90007, USA.
The Journal of the Acoustical Society of America
|January 3, 2020
Summary
This study introduces a modular architecture for articulatory speech synthesis, enabling realistic vocal tract and glottis modeling for improved speech production research.
Area of Science:
- Speech Synthesis
- Articulatory Phonetics
- Biomedical Engineering
Background:
- Articulatory speech synthesis requires integrated models of the vocal tract, glottis, and aerodynamics.
- Previous models often lack modularity and integration of real-time imaging data.
Purpose of the Study:
- To propose a novel modular architecture for articulatory speech synthesis.
- To integrate diverse models for vocal tract, glottis, aero-acoustics, and control.
- To validate the architecture using speaker-specific data and synthesis tasks.
Main Methods:
- Developed a modular architecture combining statistical articulatory models (from real-time MRI data) with aerodynamic and glottal models (Maeda's work).
- Implemented an articulatory control module using dynamical systems inspired by task dynamics.
- Utilized an alpha-beta model for midsagittal to area function conversion.
Main Results:
- Successfully synthesized vowel-consonant-vowel sequences with plosive consonants.
- Demonstrated speaker-specific modeling by building and simulating models for two individuals.
- Validated the modular approach for articulatory speech synthesis.
Conclusions:
- The proposed modular architecture provides a flexible and effective framework for articulatory speech synthesis.
- Integration of statistical articulatory models with aerodynamic and control systems enhances synthesis realism.
- The approach is adaptable for simulating different speakers and their articulatory behaviors.
More Related Videos
Related Concept Videos
Elaborative Rehearsals
274
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
274
Components of Language
689
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
689
Multi-input and Multi-variable systems
344
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
344
Modeling and Similitude
555
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
555
State Space Representation
478
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
478
Design Example
497
The innovation of touch-tone telephony revolutionized the telecommunications industry by replacing the traditional rotary dial with a dual-tone multi-frequency (DTMF) signaling system. This system uses a matrix-style keypad with buttons arranged in four rows and three columns, creating 12 distinct signals each assigned to a pair of frequencies. Each button press results in a simultaneous generation of two sinusoidal tones – one from a low-frequency group (697 to 941 Hz) and one from a...
497

