Related Experiment Video
Updated: Oct 10, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.0K
Transformer-based CNNs: Mining Temporal Context Information for Multi-sound COVID-19 Diagnosis
Summary
This study shows computer audition can detect COVID-19 using breathing and speech sounds. Transformer-based deep learning models achieved 70% accuracy, offering a fast, low-cost screening method.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Bioacoustics
Background:
- Early COVID-19 screening is crucial for transmission control.
- Computer audition presents a rapid, cost-effective, and eco-friendly diagnostic approach.
- Respiratory and speech sounds contain valuable clinical information for COVID-19 detection.
Purpose of the Study:
- To develop and evaluate deep neural networks for COVID-19 detection using respiratory and speech sounds.
- To improve diagnostic performance by assembling multiple models trained on different sound types.
- To leverage advanced deep learning architectures, including CNNs and transformers, for enhanced feature extraction and temporal context mining.
Main Methods:
- Training three deep neural networks on breathing, counting, and vowel sounds.
- Employing Convolutional Neural Networks (CNNs) for spatial feature extraction from log Mel spectrograms.
- Utilizing a multi-head attention mechanism within a transformer to capture temporal dependencies.
Main Results:
- Transformer-based CNNs demonstrated effective COVID-19 detection on the DiCOVA Track-2 database.
- Achieved an Area Under the Curve (AUC) of 70.0%.
- Outperformed traditional CNNs and hybrid CNN-Recurrent Neural Networks (RNNs).
Conclusions:
- Deep learning models, particularly transformer-based CNNs, show significant potential for COVID-19 detection via computer audition.
- Assembling models trained on diverse sound types enhances diagnostic accuracy.
- This approach offers a promising avenue for accessible and efficient COVID-19 screening.

