Related Experiment Video
Updated: Aug 5, 2026

12:43
A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
Published on: February 21, 2011
Robust nasality representation learning for cleft palate-related velopharyngeal dysfunction screening in real-world
Weixin Liu1, Bowen Qu2, Amy Stone3
1Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, United States.
Frontiers in Digital Health
|July 31, 2026
Summary
A new machine learning framework improves velopharyngeal dysfunction (VPD) screening by learning nasality representations, enhancing accuracy in real-world conditions. This approach boosts performance on diverse recordings, overcoming limitations of clinical settings.
Area of Science:
- Speech pathology
- Machine learning
- Digital health
Background:
- Velopharyngeal dysfunction (VPD) impairs speech, causing hypernasality and reduced intelligibility.
- Current screening requires specialized settings, limiting accessibility globally.
- Machine learning models struggle with real-world audio due to domain shift from varied recording conditions.
Purpose of the Study:
- To develop a robust two-stage framework for VPD screening adaptable to uncontrolled acoustic environments.
- To enhance the reliability of speech-based diagnostic tools for wider deployment.
- To mitigate performance degradation caused by domain shift in consumer devices.
Main Methods:
- A two-stage framework incorporating nasality representation pre-training using supervised contrastive learning.
- Frozen-encoder VPD screening on short audio segments with probability aggregation for decision-making.
- Evaluation using both in-domain (clinical) and out-of-domain (internet) datasets, compared against baselines like MFCC and large speech representations.
Main Results:
- Achieved perfect screening performance (macro-F1=1.000, accuracy=1.000) on in-domain clinical data.
- Demonstrated robust performance on out-of-domain internet recordings (macro-F1=0.679, accuracy=0.695), outperforming strong baselines.
- The proposed method showed significant improvements over MFCC and large pretrained speech representations on heterogeneous recordings.
Conclusions:
- Pre-training with nasality-focused representations improves robustness against recording artifacts.
- The framework supports practical deployment of VPD screening tools in real-world scenarios.
- Highlights the need for domain-robust evaluation protocols for speech-based digital health technologies.
Related Concept Videos
Suctioning the Nasopharyngeal Airway
Nasopharyngeal suctioning is a procedure to remove secretions from the upper part of the respiratory tract that the patient cannot clear independently. It helps maintain airway patency and prevents complications such as aspiration pneumonia.
Equipment Required
Equipment Required
Cardiopulmonary Resuscitation II: ACLS Airway Management
Airway management is a key skill in emergency and critical care settings, as maintaining a clear airway is essential for adequate oxygenation and ventilation.Head Tilt-Chin Lift TechniqueThe head tilt-chin lift maneuver is an essential technique primarily used in patients without suspected cervical spine injuries. To perform this maneuver, one hand is placed on the patient’s forehead, and gentle pressure is applied backward to tilt the head. The fingertips of the other hand are positioned under...