Related Experiment Videos
A novel class-attention transformer-driven feature fusion technique-based speech disorder classification
Abdul Rahaman Wahab Sait1, Haitham Ahmed Jamil Mohammed2, Taqwa Ali Mohammad Bani Awad3,4
1Department of Documents and Archive, Center of Documents and Administrative Communication, King Faisal University, Al-Ahsa, Saudi Arabia.
Frontiers in Medicine
|July 8, 2026
Summary
This study introduces a novel framework for detecting speech disorders (SD) using a hybrid CNN-ViT model. The approach significantly improves diagnostic accuracy and interpretability for speech pathology applications.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Speech Pathology
Background:
- Speech disorders (SD) pose diagnostic challenges due to complex acoustic features in speech.
- Existing methods using standalone CNNs and ViTs struggle to capture the temporal and spectral dynamics of SD.
Purpose of the Study:
- To develop an end-to-end framework for binary pathological SD detection using raw audio waveforms.
- To improve feature representation, interpretability, and generalization in SD detection.
Main Methods:
- A hybrid feature extraction method combining 1D CNNs and ViT's self-attention mechanism.
- Adaptive fusion and Class-attention transformer (CaiT)-based refinement for classification-optimized representation.
- Integration of Grad-CAM and attention-based visualization for temporal localization of disorder-relevant patterns.
Main Results:
- Achieved 97.50% accuracy on SVD and PD datasets, and 95.80% on VOICED dataset.
- The model is lightweight, featuring only 5.2 million parameters.
- Statistical analysis confirmed the reliability and significance of the results.
Conclusions:
- The hybrid framework enhances SD detection performance and interpretability.
- Offers a proof-of-concept for reliable SD diagnosis, aiding clinicians and speech-language pathologists.
- Enables focused attention on diagnostically significant time intervals within speech signals.