Related Experiment Video
Updated: Apr 21, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
15.2K
Sparse Clustering with Resampling for Subject Classification in PET Amyloid Imaging Studies.
Wenzhu Bi1, George C Tseng2, Lisa A Weissfeld2
1Department of Biostatistics, University of Pittsburgh, Pittsburgh PA 15261 USA telephone: 412-605-1552.
Summary
Sparse k-means clustering with data resampling effectively identifies informative variables and establishes confidence levels for clustering, particularly in high-dimensional data. This method successfully distinguished amyloid-positive from amyloid-negative subjects in PiB PET imaging.
Area of Science:
- Computational biology and bioinformatics
- Neuroimaging analysis
- Statistical machine learning
Background:
- Sparse k-means clustering (Sparse_kM) is valuable for parsimonious clustering in high-dimensional datasets (p≫n).
- Identifying informative variables and defining cluster confidence are crucial for reliable analysis.
Purpose of the Study:
- To combine Sparse k-means clustering with data resampling to identify key variables and establish confidence levels.
- To apply this enhanced method to PiB PET imaging data for classifying normal control subjects based on amyloid status (PiB+/-).
Main Methods:
- Statistical simulations using a dataset with n=60 observations and p=500 variables, where only 50 variables were truly informative.
- Data resampling (20 times) followed by Sparse_kM application to calculate average variable weights and cluster membership probabilities (confidence levels).
- Application to a PiB PET dataset (n=64 subjects, p=343,099 voxels) with 10 resampling iterations to identify informative voxels and subject classification.
Main Results:
- Simulations correctly identified the 50 truly different variables with weights 13-32 times greater than uninformative variables.
- The PiB PET analysis highlighted key cortical areas (precuneus, frontal cortex) associated with amyloid deposition.
- Seven subjects were classified as PiB(+) and 47 as PiB(-) with ≥90% confidence; one subject was PiB(+) with 80% confidence; nine were PiB(-) with 50-70% confidence.
Conclusions:
- Sparse k-means clustering combined with data resampling provides robust confidence levels for clustering high-dimensional data.
- This approach effectively identifies informative features, such as voxels in amyloid PET imaging, distinguishing different disease states.
- The method shows promise for analyzing complex neuroimaging data and classifying subjects, especially at transitional disease boundaries.

