Related Experiment Video
Updated: Jul 9, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Data-driven refinements for voice disorder classification: improving accuracy and generalisability
Rijul Gupta1, Catherine Madill2, Craig Jin1
1Computing and Audio Research Laboratory, School of Electrical and Computer Engineering, The University of Sydney, Sydney, NSW, Australia.
Introduction:
As machine-learning models for vocal pathology advance, their performance is increasingly constrained not by modelling techniques but by the taxonomic structures used to define the classification task itself. Conventional clinical frameworks, while grounded in diagnostic practice, often reflect conceptual groupings that do not map cleanly onto the acoustic patterns learned by modern Voice AI systems-contributing to the persistent performance gap between multi-class and binary detection tasks. Motivated by this mismatch, we introduce an alternative strategy: deriving a taxonomy from data-driven acoustic relationships rather than prescriptive clinical categories, with the goal of establishing a more model-aligned and generalisable foundation for voice disorder classification.
Methods:
We developed CarLab 2025, a novel data-driven classification framework derived from model confusion patterns. We conducted comprehensive experiments comparing its performance against existing clinical taxonomies, including the hierarchical USVAC 2025 framework, as well as Compton 2022, da Silva Moura 2024, and Za'im 2023, across multiple vocal tasks, features, and model architectures. We evaluated both in-domain performance and cross-database generalisation, including experiments with multi-task learning and targeted data injection.
Results:
CarLab 2025 achieved superior in-domain classification accuracy compared to established clinical taxonomies, with balanced accuracy reaching 67.20% compared to 61.03% for the best-performing clinical framework. For out-of-domain generalisation, models trained with structured taxonomies consistently outperformed those trained with narrow, single-disorder labels, and training on a diverse set of vocal tasks proved more effective for cross-database performance than relying on a single task. Multi-task learning offered no advantage over single-task training, and while injecting a small amount of data from target domains significantly boosted binary detection accuracy, this improvement did not consistently translate to multi-class recall.
Discussion:
Our experiments established a baseline performance exceeding that obtained with existing clinical classification frameworks by aligning more closely with acoustic manifestations of disorders. We further show that exposure to varied recording conditions is crucial for binary generalisation, while robust multi-class generalisation will require substantially more diverse multi-source training data. The results provide a clear, evidence-based path toward developing more robust and generalisable models for vocal pathology detection.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Improving Translational Accuracy
Improving Translational Accuracy