Related Experiment Video
Updated: Aug 21, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Foundations for pediatric vocal biomarkers: age-aware phoneme recognition and latent-space error analysis
Chethana Saligram1, Vishal Shrivastava1, Marisha Speights1
1Roxelyn and Richard Pepper Department of Communication Disorders, Northwestern University, Evanston, IL, United States.
Introduction:
Pediatric speech sound disorders (SSDs) affect many young children and are commonly assessed through auditory-perceptual judgments and IPA transcription, which are limited by listener bias, variable interrater reliability, and difficulty attributing deviations at the phoneme level. We present a developmentally informed framework for pediatric vocal biomarker foundations that prioritizes age-aware, phoneme-resolved interpretability over utterance-level accuracy alone.
Methods:
Using a CAAP-derived subset of the SEED corpus (27-94 months) with clinically specified phoneme targets and SSD labels, we construct age-stratified phoneme profiles and quantify error structure via phoneme error rate (PER) and age-banded confusion signatures. We operationalize co-articulation as transition dynamics from MFCC trajectories and their first and second derivatives, reflecting the velocity and acceleration of spectral change across adjacent segments. We learn compact acoustic representations with a variational autoencoder (VAE) and model temporal evolution with BiLSTMs, including attention, to characterize disorder-relevant instability in latent trajectories. For phoneme transcription, we train BiLSTM-CTC sequence models on clinically elicited speech and evaluate disorder classification and phoneme substitution patterns within age groups, then stress-test generalization on ECSC "Frog Story" narratives from TalkBank/CHILDES.
Results:
Co-articulation transition magnitudes varied systematically by age and clinical group. Attention improved temporal classification specifically in younger children, and PER dropped markedly between ages 3-4 and 4-5. On long-form naturalistic narratives, CTC training and inference remained numerically stable, but greedy decoding produced degenerate output dominated by blank and repetition tokens.
Discussion:
Together, these results support an age-aware approach to pediatric phoneme analytics, in which articulatory coordination, not just phoneme identity, carries developmentally and clinically relevant information; and phoneme-level performance should be interpreted relative to developmental stage rather than a single fixed benchmark. The decoding failures on naturalistic speech indicate a bottleneck specific to decoding rather than to the underlying representations, motivating constrained, hierarchical, or duration-aware decoding as a tractable next step. These findings position age-stratified, phoneme-resolved analysis as a foundation for interpretable and scalable pediatric screening and biomarker development.