Related Experiment Video
Updated: Sep 1, 2025

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Data augmentation using Variational Autoencoders for improvement of respiratory disease classification
Jane Saldanha1, Shaunak Chakraborty2, Shruti Patil1
1Symbiosis Centre for Applied Artificial Intelligence, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune, Maharashtra, India.
Synthesizing respiratory sounds using Variational Autoencoders (VAEs) helps overcome imbalanced datasets for improved deep learning classification. Augmenting data with synthetic sounds significantly boosts lung sound diagnostic accuracy.
Area of Science:
- Medical informatics
- Artificial intelligence in healthcare
- Respiratory medicine
Background:
- Computerized auscultation of lung sounds is crucial for diagnosing respiratory diseases, but existing datasets like ICBHI are imbalanced.
- Imbalanced datasets hinder the generalization and reliability of deep learning models for lung sound classification.
- Traditional diagnostic methods have limitations that advanced computational approaches aim to surpass.
Purpose of the Study:
- To synthesize respiratory sounds using various Variational Autoencoder (VAE) models, including Multilayer Perceptron VAE (MLP-VAE), Convolutional VAE (CVAE), and Conditional CVAE.
- To evaluate the impact of augmenting an imbalanced respiratory sound dataset with synthesized data on lung sound classification model performance.
- To compare the quality of synthesized respiratory sounds using metrics like Fréchet Audio Distance (FAD).
Main Methods:
- Respiratory sound synthesis using MLP-VAE, CVAE, and Conditional CVAE.
- Dataset augmentation with synthesized respiratory sounds.
- Performance evaluation of lung sound classification models on augmented datasets.
- Quality assessment of synthetic sounds using Fréchet Audio Distance (FAD), Cross-Correlation, and Mel Cepstral Distortion.
Main Results:
- Convolutional VAE (CVAE) and Conditional CVAE demonstrated superior synthetic sound quality with average FAD scores of 11.58 and 11.64, respectively, compared to MLP-VAE's 12.42.
- Augmenting the imbalanced dataset with synthesized sounds led to significant performance improvements in classification metrics for minority classes.
- Marginal performance gains were observed for other classes after data augmentation.
Conclusions:
- Deep learning models show promise for lung sound classification, offering advantages over traditional methods.
- Synthesizing respiratory sounds with VAEs and augmenting imbalanced datasets can significantly enhance the performance of lung sound classification models.
- The study highlights the potential of generative AI techniques in improving the accuracy and reliability of computational diagnostics for respiratory conditions.
Related Concept Videos
Respiratory Volumes and Capacities I
Respiratory Volumes
Tidal Volume (TV) Tidal volume (TV) is the air inhaled or exhaled in a...
Respiratory Volumes and Capacities
Assessment of Ventilation II: Respiratory Depth and Rhythm
Respiratory depth measures the volume of air inhaled or exhaled during a breath. It can vary from shallow to deep and typically remains consistent when a person is at rest or asleep. Occasionally, individuals will automatically inhale deeply, known as sighing, which inflates the lungs with more air than normal breathing.
To assess respiratory depth, observe the degree of chest excursion or movement:
Common Respiratory Disorders
Upper respiratory disorders impact the airways above the vocal cords, encompassing areas like the nose, sinuses, and throat. Various conditions fall under this category, including the common cold and allergic rhinitis. These disorders can stem from several causes,...
Factors Affecting Pulmonary Ventilation
Alveolar Surface Tension
The alveolar fluid lines the luminal surface of the alveoli and exerts a force called surface tension. This force is caused by the polar water molecules in the liquid being more strongly attracted to each...

