Related Experiment Video
Updated: Apr 11, 2026

11:15
fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
7.2K
The Mason-Alberta Phonetic Segmenter: a forced alignment system based on deep neural networks and interpolation
Matthew C Kelley1, Scott James Perry2, Benjamin V Tucker2,3
1Department of English, Linguistics Program, George Mason University 3298 , Fairfax, VA, USA.
Phonetica
|September 9, 2024
Summary
We developed a new neural network forced alignment system, Mason-Alberta Phonetic Segmenter (MAPS), that significantly improves speech segment boundary detection. MAPS outperforms current state-of-the-art systems, offering more precise results for large speech corpora analysis.
Area of Science:
- Computational Linguistics
- Speech Processing
- Artificial Intelligence
Background:
- Forced alignment systems are crucial for segmenting speech data, enabling large-scale corpus analysis.
- Existing systems often have limitations in boundary precision and acoustic modeling approaches.
Purpose of the Study:
- To introduce the Mason-Alberta Phonetic Segmenter (MAPS), a novel neural network-based forced alignment system.
- To evaluate two key improvements: treating acoustic models as taggers and employing an interpolation technique for precise boundaries.
Main Methods:
- Developed a neural network architecture for forced alignment.
- Implemented an acoustic model as a tagger, acknowledging overlapping speech segments.
- Utilized an interpolation technique to refine boundary detection beyond the typical 10 ms limit.
Main Results:
- All MAPS configurations significantly outperformed the Montreal Forced Aligner at a 10 ms tolerance.
- Achieved up to a 28.13% relative performance increase in boundary placement accuracy.
- Observed the Montreal Forced Aligner slightly outperforming MAPS only at a 30 ms tolerance.
Conclusions:
- MAPS demonstrates superior performance in precise phonetic boundary detection compared to state-of-the-art systems.
- The study highlights potential discrepancies between acoustic model training targets and phonetic similarity, suggesting future research directions.
- Rethinking acoustic modeling and output targets may be necessary for further advancements in forced alignment.
Keywords:
acoustic modelsautomatic speech recognitionforced alignmentneural networksphoneticsspeech technologyMore Related Videos
Related Concept Videos
One-Degree-of-Freedom System
954
In mechanical engineering, one-degree-of-freedom systems form the basis of a wide range of electrical and mechanical components. Using these models, engineers can predict the behavior of various parts in a larger system, which gives them insight into how different forces interact with each other.
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
954
Reconstruction of Signal using Interpolation
878
Signal processing techniques are essential for accurately converting continuous signals to digital formats and vice versa. When a continuous signal is sampled with a period T, the resulting sampled signal exhibits replicas of the original spectrum in the frequency domain, spaced at intervals equal to the sampling frequency. To handle this sampled signal, a zero-order hold method can be applied, which creates a piecewise constant signal by retaining each sample's value until the next...
878

