Related Experiment Video
Updated: Aug 17, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Optimal Acoustic Parameter Subsets for Dimension-Specific Voice Quality Prediction
1Dept. Computer Science, University of Brasília, Brasília, DF, 70910-900, Brazil.
Objectives:
To determine the minimal set of acoustic parameters that optimally predicts each major perceptual voice quality dimension and to quantify the redundancy among commonly measured acoustic voice parameters.
Study Design:
Secondary analysis of a publicly available voice database.
Methods:
Fourteen acoustic parameters were extracted from sustained vowels of 284 speakers in the Perceptual Voice Qualities Database (PVQD). Four right-skewed parameters (jitter, shimmer, shimmer dB, and period standard deviation) were log-transformed to linearize their relationship with perceptual ratings. QR factorization with column pivoting characterized the redundancy structure among parameters. Orthogonal Matching Pursuit (OMP) ranked parameters by incremental predictive value for each perceptual dimension (CAPE-V Severity, Breathiness, Roughness, and Strain; GRBAS Grade, Breathiness, and Roughness). The optimal number of parameters was determined by convergence of adjusted R2, Bayesian Information Criterion, and cross-validated prediction error. Classification performance against dichotomized perceptual ratings was evaluated using ROC analysis with 10-fold stratified cross-validation.
Results:
The 14-parameter acoustic space had an effective dimensionality of approximately 11, with slope contributing negligible independent information. Log transformation improved predictive accuracy for all dimensions, except breathiness, with gains that were modest for overall severity and substantial for roughness. Optimal parameter subsets varied across perceptual dimensions: overall severity required six parameters (log shimmer dB, GNE, Hno-6000, log PSD, F0, and HNR-D; rs = 0.77, AUC = 0.870), breathiness required three (CPPS, GNE, and Hno-6000; rs = 0.73, AUC = 0.865), roughness required 5 (log shimmer dB, H1-H2, GNE, log PSD, and log jitter; rs = 0.71, AUC = 0.830), and strain required six (HNR, F0, log PSD, tilt, H1-H2, and GNE; rs = 0.66, AUC = 0.790). The severity subset yielded a higher area under the curve (AUC) than the Acoustic Voice Quality Index (AVQI: six parameters, AUC = 0.827) and the Cepstral Spectral Index of Dysphonia (CSID: AUC = 0.787). This advantage was obtained from sustained vowels alone, whereas AVQI and CSID were scored on connected speech in addition to the vowel; the comparison is therefore task-asymmetric and favors the established indices in speech material. Notably, log shimmer dB displaced smoothed cepstral peak prominence (CPPS) as the most important predictor of overall severity, suggesting that the widely reported primacy of CPPS may partly reflect its distributional advantage in linear models.
Conclusions:
Different perceptual dimensions of voice quality are optimally predicted by distinct subsets of acoustic parameters. Log-transforming right-skewed perturbation measures improves linear prediction incrementally for overall severity and substantially for roughness. Using sustained vowels alone, data-driven parameter selection matches or exceeds the classification performance of established multiparametric indices scored on more speech material, and identifies candidate parameter sets for roughness and strain-dimensions lacking published composite indices. External validation and extension to connected speech remain necessary before clinical adoption.
Related Concept Videos
Expected Frequencies in Goodness-of-Fit Tests
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length, the...
Sampling Methods: Overview
In analytical chemistry, the choice of sampling...