用SAM 2.0实现实时声道图像细分的自动化和形态操作实施
Haley Hsu1, Kyle Ng1, Ameen Qureshi1
1Department of Linguistics, University of Southern California, Los Angeles, California 90007, USA.
JASA express letters
|March 4, 2026
概括
细分任何东西模型2在实时MRI数据中高效地细分语音表达器. 这通过简化复杂声道动态的分析来推进语音生产研究.
科学领域:
- 语音科学 语言科学
- 生物医学成像技术 生物医学成像技术
- 计算语言学 计算语言学
背景情况:
- 准确的模拟表达表达对于理解语音产生及其声学相关的东西至关重要.
- 在连续演讲中分辨表达力学动态,会带来重大的计算挑战.
- 现有的细分方法,如实时声道成像的轮跟踪,通常需要手动模板创建和人类监督,限制效率.
研究的目的:
- 调查Segment Anything Model 2 (SAM 2.0) 的有效性,用于在实时磁共振成像 (MRI) 语音生成数据中细分关键关器.
- 评估模型在没有特定任务微调的情况下执行细分的能力.
- 评估全球非线性图像过对语言和主题特定特征的语音动态细分的影响.
主要方法:
- 用途分段任何模型2 (SAM 2.0) 用于关器的细分.
- 将模型应用于实时MRI语音生成数据.
- 集成的全球非线性图像过作为处理管道的一部分.
主要成果:
- 使用SAM 2.0在不需要微调的情况下证明了关键关节器的高效细分.
- 成功地将该模型应用于复杂的实时MRI语音数据.
- 展示了该模型在细分具有固有的变异性的动态语音特征方面的能力.
结论:
- SAM 2.0提供了一种高效有效的方法,用于在实时MRI语音数据中对发音器进行细分.
- 没有微调的模型的性能简化了语音生产动态的分析.
- 这种方法有望促进语言科学,声学和计算语言学的研究.
相关概念视频
Sampling Theorem
In signal processing, the analysis of continuous-time signals, denoted as x(t), often involves sampling techniques to convert these signals into discrete-time signals. This process is essential for digital representation and manipulation. A critical component in sampling is the train of impulses, characterized by the sampling interval and the sampling frequency. The relationship between these parameters and the original signal's properties dictates the success of the sampling process.
Downsampling
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...


