Related Experiment Video
Updated: Aug 10, 2026

12:08
From Voxels to Knowledge: A Practical Guide to the Segmentation of Complex Electron Microscopy 3D-Data
Published on: August 13, 2014
25.1K
Automation of real-time vocal tract image segmentation with SAM 2.0 and morphological operation implementation
Haley Hsu1, Kyle Ng1, Ameen Qureshi1
1Department of Linguistics, University of Southern California, Los Angeles, California 90007, USA.
JASA Express Letters
|March 4, 2026
Summary
Segment Anything Model 2 efficiently segments speech articulators in real-time MRI data. This advances speech production research by simplifying the analysis of complex vocal tract dynamics.
Area of Science:
- Speech Science
- Biomedical Imaging
- Computational Linguistics
Background:
- Accurate modeling of articulatory representations is crucial for understanding speech production and its acoustic correlates.
- Discretizing articulatory dynamics in continuous speech presents significant computational challenges.
- Existing segmentation methods, like contour tracking for real-time vocal tract imaging, often require manual template creation and human supervision, limiting efficiency.
Purpose of the Study:
- To investigate the efficacy of Segment Anything Model 2 (SAM 2.0) for segmenting critical articulators in real-time magnetic resonance imaging (MRI) speech production data.
- To assess the model's ability to perform segmentation without task-specific fine-tuning.
- To evaluate the impact of global nonlinear image filtering on segmenting speech dynamics with language- and subject-specific characteristics.
Main Methods:
- Utilized Segment Anything Model 2 (SAM 2.0) for the segmentation of articulators.
- Applied the model to real-time MRI speech production data.
- Incorporated global nonlinear image filtering as part of the processing pipeline.
Main Results:
- Demonstrated efficient segmentation of critical articulators using SAM 2.0 without requiring fine-tuning.
- Successfully applied the model to complex real-time MRI speech data.
- Showcased the model's capability in segmenting dynamic speech features with inherent variability.
Conclusions:
- SAM 2.0 offers an efficient and effective approach for segmenting articulators in real-time MRI speech data.
- The model's performance without fine-tuning simplifies the analysis of speech production dynamics.
- This method holds promise for advancing research in speech science, acoustics, and computational linguistics.
Related Concept Videos
Sampling Theorem
In signal processing, the analysis of continuous-time signals, denoted as x(t), often involves sampling techniques to convert these signals into discrete-time signals. This process is essential for digital representation and manipulation. A critical component in sampling is the train of impulses, characterized by the sampling interval and the sampling frequency. The relationship between these parameters and the original signal's properties dictates the success of the sampling process.
Downsampling
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...

