Automatic detection of pathological voices using complexity measures, noise parameters, and mel-cepstral coefficients
Julián D Arias-Londoño1, Juan I Godino-Llorente, Nicolás Sáenz-Lechón
1Department ICS, Universidad Politecnicade Madrid, Madrid 28031, Spain. jdariasl@unal.edu.co
IEEE Transactions on Bio-Medical Engineering
|January 25, 2011
Summary
This study enhances pathological voice detection by combining nonlinear speech analysis with traditional methods. The novel approach achieved a high accuracy of 98.23% for identifying voice disorders.
Area of Science:
- Biomedical Engineering
- Signal Processing
- Speech Science
Background:
- Accurate detection of pathological voices is crucial for diagnosis and treatment.
- Existing methods often rely on conventional acoustic features, which may not capture all relevant voice characteristics.
- Nonlinear analysis offers potential for extracting more informative features from speech signals.
Purpose of the Study:
- To improve the accuracy of automatic pathological voice detection systems.
- To explore the utility of nonlinear time series analysis features for voice pathology discrimination.
- To develop a robust classification strategy by fusing nonlinear features with traditional acoustic parameters.
Main Methods:
- Extraction of 11 features using nonlinear analysis: largest Lyapunov exponent, correlation dimension, recurrence, fractal-scaling, and entropy estimations.
- Integration of nonlinear features with conventional parameters like noise measures and mel-frequency cepstral coefficients (MFCCs).
- A two-step classification approach employing a generative model followed by a discriminative model.
Main Results:
- The combined classifier achieved a high accuracy of 98.23% ± 0.001 in detecting pathological voices.
- Nonlinear features demonstrated significant discrimination capabilities when fused with traditional parameters.
- The proposed two-step classification strategy effectively leveraged the strengths of both generative and discriminative models.
Conclusions:
- The proposed approach significantly enhances the accuracy of pathological voice detection.
- Combining nonlinear speech analysis with classic parameterization offers a powerful strategy for voice disorder identification.
- This method holds promise for improving diagnostic tools in speech pathology.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Classification of Signals
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
