Related Experiment Video
Updated: Jul 18, 2025

08:05
Lung CT Segmentation to Identify Consolidations and Ground Glass Areas for Quantitative Assesment of SARS-CoV Pneumonia
Published on: December 19, 2020
14.2K
A semi-supervised algorithm for improving the consistency of crowdsourced datasets: The COVID-19 case study on
Lara Orlandic1, Tomas Teijeiro2, David Atienza1
1Embedded Systems Laboratory (ESL), EPFL, Lausanne, Switzerland.
Computer Methods and Programs in Biomedicine
|August 20, 2023
Summary
Semi-supervised learning (SSL) enhances cough audio classification by improving data consistency for COVID-19 detection and cough characterization. This method aggregates expert knowledge, creating a more reliable dataset for training diagnostic models.
Area of Science:
- Medical acoustics and signal processing
- Machine learning for healthcare
- Respiratory disease diagnostics
Background:
- Cough audio classification shows promise for screening respiratory disorders like COVID-19.
- Crowdsourcing cough data, like the COUGHVID dataset, is vital due to the risks of collecting data from contagious patients.
- Expert annotations in datasets can suffer from mislabeling and inter-expert disagreement.
Purpose of the Study:
- To improve the labeling consistency of the COUGHVID dataset using semi-supervised learning (SSL).
- To enhance classification accuracy for COVID-19 versus healthy coughs, cough type (wet/dry), and severity.
- To generate a reliable, augmented dataset for training cough classifiers.
Main Methods:
- Applied SSL expert knowledge aggregation techniques to address label inconsistencies and sparsity in the COUGHVID dataset.
- Utilized audio signal processing and interpretable machine learning models.
- Identified a subsample of re-labeled audio samples for training or augmenting cough classifiers.
Main Results:
- Re-labeled data showed significantly higher inter-class feature separability (3x for COVID-19 vs. healthy, 11.3x for type, 5.1x for severity).
- Amplified spectral differences in re-labeled data led to distinct power spectral densities between healthy and COVID-19 coughs (1-1.5 kHz range).
- A COVID-19 classifier trained on the re-labeled dataset achieved an AUC of 0.797.
Conclusions:
- Introduced a novel SSL expert knowledge aggregation technique for cough sound classification.
- Demonstrated an explainable method to combine multi-expert medical knowledge, yielding consistent and abundant data.
- The re-labeled dataset provides a robust foundation for improved cough classification tasks.
Related Concept Videos
Classification of Illness
7.6K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.6K
Common Respiratory Disorders
667
Respiratory disorders, a prevalent health concern globally, are generally divided into two primary categories: upper and lower respiratory tract disorders. The categorization is based on the area of the respiratory system they affect.
Upper respiratory disorders impact the airways above the vocal cords, encompassing areas like the nose, sinuses, and throat. Various conditions fall under this category, including the common cold and allergic rhinitis. These disorders can stem from several causes,...
Upper respiratory disorders impact the airways above the vocal cords, encompassing areas like the nose, sinuses, and throat. Various conditions fall under this category, including the common cold and allergic rhinitis. These disorders can stem from several causes,...
667
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K

