Synthesizing vocal tract magnetic resonance imaging sequences with phoneme-aware diffusion models

Paula Andrea Perez-Toro1,2, Tomas Arias-Vergara1,2, Lukas Buess1

  • 1Friedrich-Alexander-Universität Erlangen-Nürnberg, Pattern Recognition Lab, Erlangen, Germany.

Summary

This study introduces a diffusion-based framework for generating vocal tract MRI from speech, improving phoneme-level accuracy. The novel Speech-to-Image Phonemic Fidelity Score (SIPFS) enhances articulatory precision for speech science and clinical applications.