Comparison of machine learning methods for classifying aphasic and non-aphasic speakers.
Antti Järvelin1, Martti Juhola
1University of Tampere, Department of Information Studies and Interactive Media, FI-33014 University of Tampere, Finland. ajarvelin@gmail.com
Computer Methods and Programs in Biomedicine
|April 19, 2011
Summary
Machine learning classifiers show varied success in classifying aphasia. No single classifier excels across all datasets, emphasizing the need for dataset-specific selection for accurate aphasia classification.
Area of Science:
- Neuroscience
- Computer Science
- Artificial Intelligence
Background:
- Aphasia classification is crucial for diagnosis and treatment.
- Machine learning offers potential tools for analyzing complex language data in aphasia.
Purpose of the Study:
- To evaluate the performance of eight machine learning classifiers on three distinct aphasia-related classification tasks.
- To determine if a universally optimal classifier exists for aphasia data analysis.
Main Methods:
- Comparison of eight machine learning classifiers using three datasets: Philadelphia Naming Test data, Finnish Boston Naming Test data (Alzheimer's vs. vascular disease), and Aachen Aphasia Test data.
- Artificial data generation for smaller datasets to augment original naming data.
- Experimental testing of classifiers on confrontation naming and aphasia syndrome data.
Main Results:
- Classifiers performed successfully on the first (aphasic vs. non-aphasic) and third (aphasic syndromes) datasets.
- Less encouraging results were observed with the second dataset (Alzheimer's vs. vascular disease).
- No single machine learning classifier demonstrated superior performance across all tested aphasia datasets.
Conclusions:
- Machine learning classifiers can be effective for certain aphasia classification tasks.
- The choice of classifier should be tailored to the specific characteristics of the aphasia dataset being analyzed.
- Further research may be needed to optimize classifier performance for challenging datasets like distinguishing between Alzheimer's and vascular dementia based on naming data.

