Related Experiment Video
Updated: Jul 4, 2025

Ultrasound Images of the Tongue: A Tutorial for Assessment and Remediation of Speech Sound Errors
Published on: January 3, 2017
Evaluating acoustic representations and normalization for rhoticity classification in children with speech sound
Nina R Benway1,2, Jonathan L Preston1, Asif Salekin3
1Communication Sciences & Disorders, Syracuse University, Syracuse, New York 13244, USA.
Age-and-sex normalization significantly improved classifiers for predicting children's rhotic /ɹ/ perception. Clinically interpretable formants performed comparably to Mel frequency cepstral coefficients (MFCCs) in this speech analysis.
Area of Science:
- Speech-language pathology
- Acoustic phonetics
- Machine learning in speech science
Background:
- Accurate classification of children's speech sound production is crucial for early identification of speech sound disorders.
- Acoustic analysis offers objective measures for speech sound characterization.
- Understanding the impact of acoustic features and normalization techniques is key to developing robust speech analysis tools.
Purpose of the Study:
- To compare the effectiveness of different acoustic representations and normalization methods for classifying children's production of the rhotic sound /ɹ/.
- To evaluate the performance of formants and Mel frequency cepstral coefficients (MFCCs) in predicting rhotic versus derhotic /ɹ/.
- To assess the impact of utterance-level versus age-and-sex-based normalization on classifier accuracy.
Main Methods:
- Utilized acoustic features, including formants and MFCCs, from 350 child speakers.
- Applied z-standardization normalization, comparing within-utterance and age-and-sex-matched typical /ɹ/ data.
- Employed statistical modeling and deep neural networks for classification, with Shapley additive explanations for feature importance analysis.
Main Results:
- Age-and-sex normalization significantly enhanced classifier performance compared to utterance-level normalization.
- Clinically interpretable formants achieved performance comparable to MFCCs.
- Personalized deep neural network models reached a mean F1-score of 0.81 (SD=0.10) for participant-specific prediction.
- Shapley additive explanations identified the third formant as the most influential feature for predicting fully rhotic productions.
Conclusions:
- Age-and-sex normalization is a critical factor for improving the accuracy of automated speech analysis in children.
- Formant-based acoustic features are viable alternatives to MFCCs for deep learning applications in speech perception research.
- The findings support the use of personalized, data-driven models for accurate assessment of phonological development, with formants offering clinical interpretability.
More Related Videos
06:04Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024