Related Experiment Video
Updated: Jul 4, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Feature Reduction or Sample Reduction? A Stability Analysis of Parkinson's Disease Clustering.
1Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School, Hannover Medical School, Hannover, Germany.
Clustering Parkinson's disease (PD) subtypes is unstable. Reducing patient samples significantly impacts results more than reducing data features, suggesting cohort size is key for reliable PD heterogeneity analysis.
Area of Science:
- Neuroscience
- Computational Biology
- Biostatistics
Background:
- Clustering methods are vital for understanding Parkinson's disease (PD) phenotypic heterogeneity.
- However, the reproducibility of PD subtype solutions derived from clustering remains a significant challenge.
Purpose of the Study:
- To investigate the impact of feature and sample reduction on clustering stability in Parkinson's disease datasets.
- To determine whether reducing data features or patient samples has a greater effect on the reliability of PD subtype identification.
Main Methods:
- Utilized baseline data from the Parkinson's Progression Markers Initiative (PPMI) cohort.
- Applied K-means, Gaussian Mixture Models (GMM), and DBSCAN algorithms.
- Systematically reduced features and samples (40%, 60%, 80%, 100%) and assessed cluster stability using Adjusted Rand Index (ARI).
Main Results:
- Sample reduction demonstrated a more pronounced effect on agreement with reference solutions across all clustering methods compared to feature reduction.
- Feature reduction primarily impacted the run-to-run reproducibility of clustering solutions.
- K-means exhibited the highest robustness, while GMM and DBSCAN showed reduced reproducibility under feature reduction and altered noise assignment, respectively.
Conclusions:
- Cohort size, reflected by sample reduction, appears to be a more critical factor for reference-solution stability in PD clustering than the number of features.
- Findings highlight the sensitivity of clustering stability to data subsetting in PD research.
- Emphasizes the need for careful consideration of sample size and feature selection in studies of PD heterogeneity.
More Related Videos
10:28Dynamic Digital Biomarkers of Motor and Cognitive Function in Parkinson's Disease
Published on: July 24, 2019
07:26Characterizing the Relationship Between Eye Movement Parameters and Cognitive Functions in Non-demented Parkinson's Disease Patients with Eye Tracking
Published on: September 26, 2019
Related Concept Videos
Parkinson Disease ll: Pathophysiology
Parkinson Disease l: Introduction
Parkinson's Disease: Treatment
Parkinson's Disease is primarily a result of the loss of dopaminergic neurons in the substantia nigra pars compacta. The cornerstone of its...
Parkinson's Disease: Overview
Neural Regulation
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...