Minimal Perturbation-Based Segment Selection in Connected Speech Differentiates Healthy and Pathological Voice
Owen P Wischhoff1, Maiwand M Tarazi1, Jakob R Holm1
1Department of Otolaryngology-Head and Neck Surgery, University of Wisconsin School of Medicine and Public Health, Madison, Wisconsin.
Objective:
Acoustic voice analysis has historically relied on sustained vowel phonation, which captures a single, controlled phonatory configuration and may underrepresent everyday communicative function. The minimal perturbation method proposes that the most regularly phonated segment of a voice sample reflects the functional capacity of the laryngeal system. This study applied a moving-window segmentation technique to connected speech to evaluate nonlinear dynamic, perturbation, and spectral acoustic measures for discriminating between healthy and pathological voices.
Methods:
Eighty-two connected speech recordings (41 healthy, 41 pathological) were randomly selected from the KayPENTAX Disordered Voice Database. The first twelve seconds of the Rainbow Passage were segmented into 0.8-second windows with a 0.25-second shift. Each metric produced a minimally perturbed segment to represent that sample. Acoustic measures included nonlinear energy difference ratio (NEDR), weighted index of dysphonia (WID), voice-type component profile (VTCP), perturbation (jitter and shimmer), cepstral peak prominence (CPP), and signal-to-noise ratio (SNR). Group differences were evaluated using Welch's independent-samples t tests with Cohen's d. Multiple comparisons were controlled using the Benjamini-Hochberg procedure within metric groups.
Results:
Seven of ten acoustic measures demonstrated significant group differences that survived within-group false discovery rate correction. CPP produced the largest effect (d = 1.953), followed by VTC1 (d = 0.711), WID (d = 0.629), NEDR (d = 0.569), SNR (d = 0.568), jitter (d = 0.530), and shimmer (d = 0.490). VTC4 reached nominal significance but did not survive correction; VTC2 and VTC3 did not reach significance.
Conclusion:
The moving-window method successfully classified healthy and pathological voices without using sustained vowels. Spectral and nonlinear dynamic measures produced the strongest effects, while segment-level optimization preserved the discriminative validity of perturbation measures in running speech. These findings support the minimal perturbation method in connected speech using these acoustic variables. Future research should further investigate optimal window approaches for running speech, focusing on composite acoustic analysis assessment to determine the lowest perturbation, balancing signal quality with fidelity.
