Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Quantum Approaches for Dysphonia Assessment in Small Speech Datasets
Abstract:
Dysphonia, a prevalent medical condition, leads to voice loss, hoarseness, or speech interruptions. To assess it, researchers have been investigating various machine learning techniques alongside traditional medical assessments. Convolutional Neural Networks (CNNs) have gained popularity for their success in audio classification and speech recognition. However, the limited availability of speech data, poses a challenge for CNNs. This study compares the performance of CNNs with standard convolutional layers, CNNs with random nonlinear convolutional layers (RANDOM models), and a novel hybrid quantum-classical architecture, Quanvolutional Neural Networks (QNNs), which are well-suited for small datasets. The audio data was preprocessed into Mel spectrograms, comprising 243 training samples and 61 testing samples in total, and used in four experiments. A total of six models were developed: two CNNs, two RANDOM models, and two QNNs, with the second models incorporating additional layers to boost performance. The angle encoding method was employed within the quanvolutional layer, which is the quantum transformation layer of the QNNs, to enable the representation of more complex features of Mel spectrogram. The experimental results demonstrated that QNN models consistently outperformed the CNN models in both classification accuracy (ranging from 82.22% ± 4% to 89.26% ± 2%) and convergence speed across all four experiments.Clinical relevance- This work investigates the potential of quantum-based approaches for medical data classification and their promising role in enhancing dysphonia assessment.

