Related Experiment Video
Updated: Jun 26, 2026

Determining 3D Flow Fields via Multi-camera Light Field Imaging
Published on: March 6, 2013
A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI
Paula Andrea Pérez-Toro1, Tomás Arias-Vergara1, Fangxu Xing2
1Harvard Medical School/Massachusetts General Hospital, Boston, 02114, Massachusetts, USA; Pattern Recognition Lab, Department of Computer Science, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, 91058, Bayern, Germany; GITA Lab, Faculty of Engineering, Universidad de Antioquia, Medellín, 050010, Antioquia, Colombia.
This study presents a novel audio-to-video framework using modified stable diffusion to generate realistic vocal tract visualizations from speech. The system aids in clinical assessment and personalized speech therapy by accurately depicting vocal tract motion.
Area of Science:
- Medical Imaging
- Speech Science
- Artificial Intelligence
Background:
- Understanding vocal tract motion during speech is vital for clinical assessment and personalized rehabilitation.
- Generating synchronized audio-visual data of the vocal tract is challenging.
Purpose of the Study:
- To develop an audio-to-video framework for generating Real Time/cine-Magnetic Resonance Imaging (RT-/cine-MRI) of the vocal tract from speech signals.
- To enable personalized simulation and outpatient therapy for speech disorders.
Main Methods:
- Temporal alignment of RT-/cine-MRI sequences and speech samples.
- Utilizing a modified stable diffusion model with structural and temporal blocks for synchronized data.
- Generating MRI sequences from novel speech inputs.
Main Results:
- The framework successfully generated realistic and accurate vocal tract visualizations from speech.
- Demonstrated adaptability to new speech inputs and effective generalization in healthy controls and tongue cancer patients.
- Positive human evaluations confirmed the quality and utility of the synthesized videos.
Conclusions:
- The developed framework effectively converts audio speech signals into visual vocal tract motion (RT-/cine-MRI).
- It shows significant potential for clinical applications, including outpatient therapy and personalized vocal tract simulations.
- This technology advances the integration of AI in speech analysis and rehabilitation.

