Related Experiment Video
Updated: May 22, 2026

fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
Synthesizing vocal tract magnetic resonance imaging sequences with phoneme-aware diffusion models
Paula Andrea Perez-Toro1,2, Tomas Arias-Vergara1,2, Lukas Buess1
1Friedrich-Alexander-Universität Erlangen-Nürnberg, Pattern Recognition Lab, Erlangen, Germany.
Purpose:
Real-time speech MRI offers critical insights into articulatory dynamics for diagnostics, therapy, and speech science, but direct speech-to-MRI mapping remains highly challenging. We present a diffusion-based framework that transforms spoken language into dynamic vocal tract MRI sequences with phonemic-aware attention. This improves phoneme-level fidelity and introduces the speech-to-image phonemic fidelity score (SIPFS), a new metric to assess articulatory precision.
Approach:
We trained audio-conditioned diffusion models with phonemic-aware attention and contrastive learning on the USC 75-Speaker Speech MRI Database. Input consisted of sagittal-view MRI sequences synchronized with speech. Three pre-trained audio encoders (Wav2Vec, WavLM, and Whisper) were used to provide embeddings. Performance was evaluated on unseen speakers and unseen speech segments using standard image/video metrics (FID, FVD, SSIM, and PSNR) and the proposed SIPFS.
Results:
Phonemic-aware attention consistently improved quality and phoneme-level precision. SemDiff produced high-quality static images, whereas SD3D excelled in dynamic video synthesis. SIPFS revealed up to 15% higher phonemic accuracy compared with unconditioned models. Excluding articulatory-region ROI reduced fidelity of vocal tract motion, highlighting its importance. The models generalized across new speakers and missing segments, demonstrating robustness.
Conclusions:
Our framework demonstrates the feasibility of synthesizing accurate, phoneme-sensitive vocal tract MRI from speech using diffusion models. By improving image fidelity and introducing SIPFS, this approach supports clinical applications in diagnostics, speech therapy, and articulatory research. Future work will focus on expanding datasets, optimizing computational efficiency, and refining evaluation metrics to further enhance clinical integration.
Related Concept Videos
Magnetic Resonance Imaging
Imaging Studies IV: Magnetic Resonance Imaging

