A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI

Paula Andrea Pérez-Toro1, Tomás Arias-Vergara1, Fangxu Xing2

  • 1Harvard Medical School/Massachusetts General Hospital, Boston, 02114, Massachusetts, USA; Pattern Recognition Lab, Department of Computer Science, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, 91058, Bayern, Germany; GITA Lab, Faculty of Engineering, Universidad de Antioquia, Medellín, 050010, Antioquia, Colombia.

Summary

This study presents a novel audio-to-video framework using modified stable diffusion to generate realistic vocal tract visualizations from speech. The system aids in clinical assessment and personalized speech therapy by accurately depicting vocal tract motion.