Related Experiment Video
Updated: Aug 5, 2026

Image Recognition and Parameter Analysis of Concrete Vibration State Based on Support Vector Machine
Published on: January 5, 2024
Sigmatism detection from child speech spectrograms using convolutional autoencoders and support vector machines
Wojciech Pieniążek1, Maria Filipek1, Oliwia Skórzewska1
1Faculty of Biomedical Engineering, Silesian University of Technology, Roosevelta 40, Zabrze, 41-800, Poland.
Background And Objective:
Sibilant consonants are among the most frequently distorted sounds in the speech of Polish children; a speech disorder involving incorrect production of these sounds is called sigmatism. The presented study investigates whether convolutional autoencoders (CAEs) can learn discriminative acoustic representations of retroflex and dental realizations of Polish sibilants and facilitate the automatic classification of sibilant place of articulation.
Methods:
A speech corpus containing recordings from 149 children producing /tʂ̑/ and /ʂ/ was segmented, normalized, and converted into spectrograms. CAE models with varying bottleneck dimensionalities (10-30) were trained, and the resulting latent vectors were classified using SVMs with radial basis function kernels under 10-fold cross-validation with separate speakers in training and test sets across folds.
Results:
For /ʂ/, the best configuration achieved 78.92% sensitivity and 86.08% accuracy, while for /tʂ̑/ sensitivity reached 74.51% with corresponding accuracy of 79.86%. The results for affricate sound classification are higher than those reported in previous works. Comparable performance across multiple latent dimensions indicates that CAEs even with smaller bottleneck layer can extract robust features from spectrograms.
Conclusions:
The results demonstrate the potential of CAE-based representations for automatic detection of non-normative sibilant articulation in children's speech, especially within affricate sounds that are rarely addressed in the literature on sigmatism detection.
