Related Experiment Video
Updated: Jun 27, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Out of the Laboratory and Into the Clinic: Out-of-Domain Validation of Machine Learning Models for Velopharyngeal
Weixin Liu1, Bowen Qu2, Amy Stone3
1Departments of Electrical Computer Engineering.
Background:
The diagnosis and management of velopharyngeal dysfunction (VPD) is a particularly challenging part of the cleft care timeline, particularly in resource-limited settings. Although machine learning (ML) models offer promising results as screening tools, their real-world clinical viability has yet to be reliably demonstrated. This study aims to systematically compare the capability of several ML models to detect VPD in nonstandardized conditions to simulate real-world clinical testing.
Methods:
Eighty-two patients were enrolled under standardized acoustic conditions and partitioned into an in-domain 60-subject training set and a 22-subject test set. For the out-of-domain test set, audio samples were obtained from publicly available Internet sources. A total of 131 case samples (70 control, 61 case) were obtained from multiple publicly available sources. Recording scenarios were highly variable, nonstandardized and largely not described, thereby introducing extreme heterogeneity.
Results:
On the hold-out testing dataset, multiple deep learning models, particularly those using Whisper and HuBERT features, achieved near-perfect performance, with the Whisper/Support Vector Machine (SVM) pipeline reaching 100% accuracy and a 1.0 macro F1-score. The traditional MFCC baseline model also performed exceptionally well, achieving 99.2% accuracy and a 0.95 macro F1-score. Cross-domain generalization testing on the out-of-domain dataset demonstrated severe performance degradation. State-of-the-art deep learning models failed to generalize. The simpler baseline MFCC/SVM pipeline proved to be the most robust in real-world simulation, achieving an accuracy of 64.1% and a macro F1-score of 0.61.
Conclusion:
The results of this study demonstrate that pretrained ML models can achieve near-perfect VPD detection in highly standardized acoustic environments. However, models are highly susceptible to domain shift caused by variations in recording devices and acoustic environments. Out-of-domain testing results were mediocre but suggest that real-world clinical detection of VPD by ML models is feasible. These results shed light on the difficulties of real-world model generalization and emphasize that for clinically deployable diagnostic tools, cross-domain robustness is as important as, if not more important than, achieving maximal accuracy on a benchmark dataset. The development of a software-based screening tool has the power to improve VPD screening, particularly in low- and middle-income countries.
More Related Videos
04:04Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025