Related Experiment Video
Updated: Jul 17, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
From voice biomarkers to telemedicine screening: developing and evaluating a voice-based AI model for laryngeal
Phillip D Jenkins1,2, Steven Bedrick1, Lisa Karstens1,3,4
1Department of Medicine, Division of Informatics and Clinical Epidemiology, Oregon Health & Science University, Portland, OR, United States.
Background:
The human voice contains rich acoustic information indicative of laryngeal pathology, yet current screening relies on resource-intensive in-person laryngoscopy. While artificial intelligence has shown promise for voice analysis, progress has been limited by small, inconsistent datasets and challenges to clinical translation. The Bridge2AI-Voice initiative addresses these barriers by providing a large-scale, ethically sourced dataset with standardized, privacy-preserving derived features.
Objective:
To determine whether the derived-feature release of Bridge2AI-Voice v3.0.0 can support a high-sensitivity screening model for laryngeal lesions and to evaluate its translational readiness using telemedicine implementation frameworks.
Methods:
We analyzed data from 205 adult participants (136 controls, 52 benign vocal fold lesions, 13 precancerous lesions, 4 laryngeal cancer) drawn from the Bridge2AI-Voice v3.0.0 derived-feature release. An L2-regularized logistic regression model was fit to 131 OpenSMILE static acoustic features with age and sex at birth, evaluated under participant-level stratified 10-fold nested cross-validation. Inner-fold cross-validation was used for operating-point threshold selection. Pre-specified validity tests against age confounding included a DeLong comparison against an age-only baseline and an age-stratified label permutation test. Alternative feature modalities (SPARC articulatory features, Mel spectrogram derivatives, and multimodal combinations) and alternative classifier families were evaluated as robustness checks.
Results:
The OpenSMILE-based model achieved cross-validated AUC 0.812 (95% CI 0.744-0.876), with operating-point sensitivity 0.870 (95% CI 0.767-0.939) and specificity 0.566 (95% CI 0.479-0.651). Model discrimination significantly exceeded an age-only baseline (DeLong p = 0.0008) and survived age-stratified label permutation (observed AUC 0.812 vs. null mean 0.553, p = 0.0099). Subgroup analysis showed approximately consistent sensitivity across benign (0.865) and precancerous (0.846) lesion subgroups. Alternative feature modalities did not provide incremental discriminative information beyond OpenSMILE, and alternative classifier families produced AUCs within bootstrap confidence intervals of the primary model.
Conclusions:
Derived acoustic features from the Bridge2AI-Voice v3.0.0 release combined with basic demographic information support cross-validated discrimination of vocal fold lesions consistent with the upper range of published voice-based laryngeal pathology classifiers. The result is presented as a candidate signal warranting confirmatory investigation in a larger, prospectively recruited cohort.