Related Experiment Videos
Spectral envelope estimation of high-pitched vowels using a shared excitation proxy and autoregressive modeling with
Zhao Zhang1, Ju Zhang2, Wenhuan Lu1
1College of Intelligence and Computing, Tianjin University, Tianjin 300350, China.
Abstract:
Characterizing the vocal tract spectral envelope in high-pitched vowels is challenging due to sparse harmonics and source-filter coupling. This study proposes an analytical framework using a shared glottal driving proxy to guide source-tract decoupling. Speech is decomposed into multi-band Hilbert envelopes; first-order differentiation and non-linear rectification are then applied to isolate transient energy increments during glottal closure. Non-negative matrix factorization captures the cross-band temporal dynamics of these increments, yielding a macroscopic temporal proxy of glottal excitation. This proxy acts as a physical constraint within an auto-regressive with exogenous input model to estimate the underlying vocal tract transfer function. To validate the estimated envelopes, formant estimation error serves as an objective metric on the OPENGLOT database. Particularly under high-pitched conditions where the fundamental frequency (F0) exceeds 300 Hz, the overall mean first formant (F1) estimation error across all evaluated vowels and parameter combinations is 11.43%, yielding an approximate 30% relative reduction from the traditional linear predictive coding baseline (16.51%). The results confirm that this approach provides a complementary pathway for vocal tract envelope modeling.