Related Experiment Video
Updated: Apr 7, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Convex weighting criteria for speaking rate estimation.
Yishan Jiao1, Visar Berisha1, Ming Tu1
1Department of Speech and Hearing Science, Arizona State University.
This study introduces a novel method for estimating speaking rate directly from speech waveforms. The approach learns an optimal weighting function, outperforming existing methods on healthy, dysarthric, and spontaneous speech.
Area of Science:
- Speech Signal Processing
- Computational Linguistics
- Machine Learning for Audio Analysis
Background:
- Estimating speaking rate directly from speech waveforms is a challenging, long-standing problem.
- Existing methods often rely on phoneme detection or envelope heuristics, which can be unreliable.
- A robust and direct method for speaking rate estimation is needed.
Purpose of the Study:
- To develop a novel method for speaking rate estimation directly from speech waveforms.
- To avoid complex phoneme detection or heuristic-based approaches.
- To learn an optimal weighting function for direct application to time-frequency features.
Main Methods:
- The problem is framed as estimating a temporal density function from speech signals.
- An optimal weighting function is learned using convex cost functions.
- An adaptation strategy is proposed for speaker-specific customization with minimal training.
- The method is evaluated on TIMIT, dysarthric, and spontaneous speech corpora.
Main Results:
- The proposed method outperforms three competing methods on healthy and dysarthric speech.
- High correlation between estimated and ground truth speaking rates is observed for spontaneous speech.
- The approach effectively estimates speaking rate without explicit phoneme detection.
Conclusions:
- The novel weighting function approach provides an effective method for speaking rate estimation.
- The technique demonstrates robustness across different speech types, including disordered speech.
- This method offers a promising alternative to traditional speaking rate estimation techniques.
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Expected Frequencies in Goodness-of-Fit Tests
Determination of Expected Frequency
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Wave Parameters
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...

