Related Experiment Video
Updated: Aug 5, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
Published on: February 21, 2011
Robust nasality representation learning for cleft palate-related velopharyngeal dysfunction screening in real-world
Weixin Liu1, Bowen Qu2, Amy Stone3
1Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, United States.
Background:
Velopharyngeal dysfunction (VPD) is an impaired ability to achieve adequate velopharyngeal closure during speech, often resulting in hypernasality and reduced intelligibility. VPD screening and diagnosis require specialized expertise and controlled recording conditions, limiting scalable access outside high-income countries.Key challenge: Speech-based machine learning models can perform extremely well under standardized clinical recording conditions. However, performance often deteriorates when deployed on consumer devices (e.g., phones or tablets) and in uncontrolled acoustic environments. This degradation is largely driven by domain shift arising from differences in recording conditions (e.g., device and channel characteristics, background noise, and room acoustics), which can cause models to rely on spurious recording artifacts rather than pathology-relevant cues.
Methods:
This study introduces a two-stage framework to improve robustness under realistic recording scenarios. Nasality representation pre-training employs a nasality-focused representation via supervised contrastive learning (SupCon) using an auxiliary dataset with phoneme alignments to form oral-context versus nasal-context supervision. During Frozen-encoder VPD screening, the encoder is frozen to perform VPD screening using lightweight classifiers on 0.5-second chunks with probability aggregation to produce recording-level decisions using a fixed decision threshold. Here, in-domain refers to standardized clinical recordings used for model development, and out-of-domain refers to heterogeneous public Internet recordings collected under uncontrolled conditions and evaluated without any adaptation. The proposed method is then compared against prior-study baselines, including MFCC features and large pretrained speech representations, using the same evaluation protocol.
Results:
On the primary in-domain subject-disjoint held-out split of 82 subjects (60 train/22 test; 345 training recordings; 131 test recordings; multiple recordings per subject), the proposed approach reached ceiling recording-level screening performance under this standardized clinical protocol (macro-F1 = 1.000, accuracy = 1.000). To assess sensitivity to this fixed split, an additional subject-level nested 5-fold cross-validation analysis was performed on the full in-domain cohort (82 subjects, 476 recordings), with the encoder frozen and only the second-stage classifiers retrained; the best mean performance was obtained with SVM (macro-F1 = 0.981 ± 0.022, accuracy = 0.985 ± 0.016). On a separate out-of-domain set of 131 public Internet recordings, large pretrained speech representations degrade substantially, and MFCC is the strongest baseline (macro-F1 = 0.612, accuracy = 0.641). The proposed method achieves the best overall out-of-domain performance (macro-F1 = 0.679, accuracy = 0.695), improving over the strongest baseline by +0.067 macro-F1 and +0.054 accuracy (point-estimate improvements) under the same evaluation protocol and fixed threshold.
Conclusion:
Learning a nasality-focused representation prior to clinical classification can reduce sensitivity to recording artifacts and improve robustness when moving from the laboratory to real-world audio recording scenarios. This design supports practical deployment of VPD screening and motivates domain-robust evaluation protocols for deployable speech-based digital health tools.
Related Concept Videos
Suctioning the Nasopharyngeal Airway
Equipment Required
Cardiopulmonary Resuscitation II: ACLS Airway Management