Related Experiment Video
Updated: Sep 5, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
V2D: Detection of type 2 diabetes using only smartphone voice recordings via spectro-temporal transformer embeddings
Abstract:
Voice-based screening offers a noninvasive, scalable avenue for early detection of type 2 diabetes using everyday smartphone recordings and acoustic features alone. We present V2D (Voice2Diabetes), a novel application of spectrogram-transformer embeddings derived exclusively from short speech segments for patient-level diabetes classification, without requiring any clinical measures or demographic variables. Adults (n=461; 157 female, 304 male) each completed multiple smartphone recordings while reading randomly selected sentences on their own smartphones. Mel-spectrograms were encoded with a pretrained Audio Spectrogram Transformer (AST) to train sex-stratified patient-level classifiers under fivefold nested cross-validation with a held-out calibration set; predictions were aggregated across recordings per participant. Using acoustic features alone, the models achieved patient-level balanced accuracy of 0.724 ± 0.023 (males) and 0.713 ± 0.021 (females), with area under the receiver operating characteristic curve (AUC) of 0.779 ± 0.019 and 0.788 ± 0.018, respectively, averaged over five independent random seeds. The pipeline incorporated probability calibration, sensitivity-first threshold optimization, and optional interpretable acoustic anchors (e.g., fundamental frequency, harmonic-to-noise ratio) to support clinical interpretation. These results provide a rigorous technical validation of AST-derived embeddings for acoustic-only, sex-stratified T2D classification in smartphone recordings and motivate prospective external validation in broader populations.