Related Experiment Video
Updated: May 22, 2026

11:15
fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
Synthesizing vocal tract magnetic resonance imaging sequences with phoneme-aware diffusion models
Paula Andrea Perez-Toro1,2, Tomas Arias-Vergara1,2, Lukas Buess1
1Friedrich-Alexander-Universität Erlangen-Nürnberg, Pattern Recognition Lab, Erlangen, Germany.
Journal of Medical Imaging (Bellingham, Wash.)
|May 21, 2026
Summary
This study introduces a diffusion-based framework for generating vocal tract MRI from speech, improving phoneme-level accuracy. The novel Speech-to-Image Phonemic Fidelity Score (SIPFS) enhances articulatory precision for speech science and clinical applications.
Area of Science:
- Medical Imaging
- Speech Science
- Artificial Intelligence
Background:
- Real-time speech Magnetic Resonance Imaging (MRI) provides crucial data for speech diagnostics, therapy, and scientific research.
- Directly mapping speech to MRI remains a significant technical challenge, limiting current applications.
Purpose of the Study:
- To develop a diffusion-based framework for synthesizing dynamic vocal tract MRI sequences directly from spoken language.
- To enhance phoneme-level fidelity in speech-to-MRI mapping.
- To introduce a new metric, the Speech-to-Image Phonemic Fidelity Score (SIPFS), for assessing articulatory precision.
Main Methods:
- Trained audio-conditioned diffusion models with phonemic-aware attention and contrastive learning on a large-scale speech MRI database.
- Utilized pre-trained audio encoders (Wav2Vec, WavLM, Whisper) for speech embedding.
- Evaluated performance using standard image/video metrics (FID, FVD, SSIM, PSNR) and the novel SIPFS on unseen speakers and speech segments.
Main Results:
- Phonemic-aware attention significantly improved image quality and phoneme-level precision.
- The proposed framework demonstrated robust generalization to new speakers and speech segments.
- SIPFS indicated up to 15% higher phonemic accuracy compared to unconditioned models, validating the approach.
Conclusions:
- The developed framework successfully synthesizes accurate, phoneme-sensitive vocal tract MRI from speech using diffusion models.
- This advancement supports clinical applications in speech diagnostics, therapy, and articulatory research.
- Future work aims to expand datasets, improve computational efficiency, and refine evaluation metrics for broader clinical integration.
Related Concept Videos
Magnetic Resonance Imaging
Magnetic resonance imaging (MRI) is a noninvasive medical imaging technique based on a phenomenon of nuclear physics discovered in the 1930s, in which matter exposed to magnetic fields and radio waves was found to emit radio signals. In 1970, a physician and researcher named Raymond Damadian noticed that malignant (cancerous) tissue gave off different signals than normal body tissue. He applied for a patent for the first MRI scanning device in clinical use by the early 1980s. The early MRI...
Imaging Studies IV: Magnetic Resonance Imaging
Introduction:Magnetic Resonance Imaging, or MRI, can include a specialized imaging technique of the urinary system known as Magnetic Resonance Urography (MRU). This radiation-free technique uses strong magnetic fields and radio waves to produce detailed images with the help of a computer. MRU is particularly effective for visualizing fluid-filled structures like the kidneys, ureters, and bladder.Applications of MRI in the Genitourinary SystemKidneys and Ureters: MRI detects tumors, cysts,...

