Related Experiment Video
Updated: Jan 16, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Fusing Time-Frequency Heterogeneous Features With Cross-Attention Mechanism for Pathological Voice Detection
Zhang Jiaqing1, Wu Yaqin1, Zhang Tao2
1Software College, Shanxi Agricultural University, Taigu 030800, China.
None:
To address the critical challenges of data scarcity, feature homogenization, and limited model generalization in current pathological voice diagnosis systems, a novel algorithm was developed to integrate time-frequency heterogeneous acoustic features for multi-class pathological voice detection. The Wav2vec2-XLSR model pretrained through self-supervised learning was first employed to extract deep contextual features from time-domain voice signals. Mel-Frequency Cepstral Coefficients (MFCC) features from the frequency domain were subsequently integrated to construct a heterogeneous vocal feature space. A cross-attention mechanism from the Transformer architecture was innovatively applied to achieve dynamic spatiotemporal alignment and semantic interaction within the heterogeneous feature space, enabling complementary feature enhancement. A dual-granularity joint analysis framework encompassing vowel and sentence hierarchies was ultimately established for efficient multi-type pathological voice detection. Experimental results demonstrated that the proposed algorithm achieved 95.1% accuracy, 100% recall, 0.92 F1-score, and 0.97 area under the ROC curve (AUC) value on sentence-level Saarbruecken Voice Database (SVD) dataset. For vowel-level classification tasks, classification accuracies of 100% and 99.6% were obtained on the Massachusetts Eye and Ear Infirmary (MEEI) and SVD datasets, respectively. Multi-corpus evaluation experiments confirm the algorithm's robustness and generalization capability across different data distributions.
More Related Videos
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

