Related Experiment Video
Updated: Jun 12, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
Published on: February 21, 2011
Characterizing Sustained Phonation in Text-To-Speech Models
Amelie Daum1, Nina Goes1, Andreas M Kist1
1Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Nürnberger Str. 74, Erlangen 91052, Bavaria, Germany.
None:
Sustained phonation (SP) is a central task in clinical voice assessment and provides a controlled setting to quantify acoustic voice characteristics. In contrast, the evaluation of modern text-to-speech (TTS) systems still relies predominantly on perceptual ratings such as the mean opinion score, leaving open whether these systems can reliably generate SP and how their acoustic properties compare to human voices. The capability of TTS models to reproduce clinically relevant voice features remains insufficiently characterized.Here, we systematically examine SP in contemporary TTS systems and compare synthetic and human voice samples using common acoustic measures. Multiple TTS models were screened for their ability to generate sustained vowels, such as /a/. One model, namely Eleven v3 by ElevenLabs, was subsequently analyzed in detail with respect to the distribution of phonation durations, the relationship between prompt length and generated duration, and differences between vowels and speaker types. Finally, TTS-generated SPs were compared with human recordings from two independent cohorts using established clinical voice parameters.We found that TTS systems were able to produce SP, although reliability varied between models. For the selected Eleven v3 model, phonation durations showed non-normal distributions and were partially predicted by prompt length. Most acoustic measures of synthetic samples overlapped with the ranges observed in human voices, while selected parameters showed statistically significant but inconsistent differences across vowels. These findings indicate that current TTS models can approximate key acoustic characteristics of SP, while also exhibiting systematic deviations that should be considered in applications involving clinical voice metrics and in further development of realistic TTS systems.
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids, corniculates, and...
Sound Waves: Resonance
Sound as Pressure Waves
The pressure fluctuation depends on the difference in displacements between the successive points in the...
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and hence a...
Echo
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case, then the...
Sound Intensity
