Interpretable Acoustic Domains in Smoking-Status Prediction From Sustained Phonation: A Speaker-Independent Secondary
Yiğit Aydoğan1, İsmail Cantürk2
1Department of Computer Science, Aberystwyth University, Aberystwyth SY23 3DB, Wales, United Kingdom.
Objectives:
To determine whether a previously developed 208-variable smoking-status voice model could be compressed into interpretable acoustic domains without materially reducing speaker-independent discrimination, and to assess which domains remained stable across model, perturbation, and counterfactual analyses.
Methods:
A secondary analysis was performed on the previously reported cohort of 64 speakers (30 smokers and 34 nonsmokers), using one primary smartphone-recorded sustained /a/ phonation per speaker. The exact 208-variable prosody-spectral feature cache was grouped into 16 prespecified acoustic domains. Domain scores were constructed independently within each leave-one-speaker-out fold using training-only imputation, robust scaling, direction alignment, and median aggregation. A balanced logistic model was evaluated from raw out-of-fold scores. Full-pipeline permutation tests, demographic residualization, domain-only models, demographic-conditional domain replacement, one-speaker jackknife analysis, a secondary Explainable Boosting Machine, and empirical counterfactual searches were performed.
Results:
The 16-domain model achieved an area under the receiver operating characteristic curve (AUC) of 0.753 (95% confidence interval [CI], 0.624-0.864), compared with 0.769 (95% CI, 0.645-0.878) for the recovered 208-variable model; the paired difference was -0.016 (95% CI, -0.129 to 0.095). Full-pipeline permutation tests were significant under both unrestricted permutation (P = 0.00599) and permutation restricted within age-by-gender strata (P = 0.00500). Spectral centroid, formant structure, and harmonicity/noise were significant as standalone domains after Holm correction. No single conditional-replacement test remained significant after correction. Spectral centroid ranked first in 49 of 64 jackknife analyses and appeared in 57.1% of minimal plausible counterfactuals. A plausible counterfactual was found for 98.4% of speakers, with one domain sufficient for 84.4%.
Conclusions:
Smoking-associated voice discrimination could be represented compactly through interpretable acoustic domains. The evidence favored a distributed, partially redundant spectral and resonance pattern rather than a single indispensable biomarker. The results describe model behavior in this cohort and require external validation before clinical interpretation.


