Linear versus nonlinear modeling of dysphonia severity: A comparison between Acoustic Voice Quality Index and machine
Ahmed M Yousef1, Adrián Castillo-Allendes2, Mark L Berardi3
1Center for Laryngeal Surgery and Voice Rehabilitation, Massachusetts General Hospital, One Bowdoin Square, 11th Floor, Boston, MA 02114, USA; Department of Surgery, Harvard Medical School, Boston, MA 02115, USA; Department of Communication Sciences and Disorders, University of Iowa, Iowa City, IA 52242, USA.
Purpose:
This study compares the Acoustic Voice Quality Index (AVQI-3) with machine learning (ML) models to evaluate their clinical utility for estimating voice quality.
Methods:
Audio from 187 American English speakers (49 healthy, 138 with voice disorders) was rated for overall voice quality by six voice specialists. AVQI-3 and its six acoustic parameters were extracted, and these parameters were used to train 14 ML models (linear, curvilinear, nonlinear). Correlations and classification accuracy (normal-mild vs. moderate-severe) were compared against perceptual ratings as ground truth.
Results:
AVQI-3 correlated strongly with perceptual ratings (Spearman rs=0.75), comparable to linear and curvilinear models, yet outperformed most nonlinear models. At a cutoff score of 2.52, AVQI-3 achieved the highest classification accuracy (0.92) with balanced sensitivity (0.90) and specificity (0.93). Among ML models, linear regression performed best (rs=0.77, accuracy=0.92, sensitivity=1.0, specificity=0.89), whereas nonlinear models showed reduced performance (average rs=0.74, accuracy=0.87, sensitivity=0.95, specificity=0.85).
Conclusion:
AVQI-3 is a simple, accessible index that quantifies voice quality as effectively as complex ML models. This is supported by the best-performing ML models being linear, indicating that linear combinations of acoustic measures are effective, accurate, and clinically interpretable, whereas added nonlinear complexity offered little benefit.
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...


