Testing for a difference in means of a single feature after clustering
1Department of Biomedical Data Science, Stanford University, 450 Serra Mall, Stanford, CA 94305, United States.
Abstract:
For many applications, it is critical to interpret and validate groups of observations obtained via clustering. A common interpretation and validation approach involves testing differences in feature means between observations in two estimated clusters. In this setting, classical hypothesis tests lead to an inflated Type I error rate. To overcome this problem, we propose a new test for the difference in means in a single feature between a pair of clusters obtained using hierarchical or k-means clustering. The test controls the selective Type I error rate in finite samples and can be efficiently computed. We further illustrate the validity and power of our proposal in simulation and demonstrate its use on single-cell RNA-sequencing data.
Related Concept Videos
Test for Homogeneity
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Significance Testing: Overview
One-Way ANOVA: Unequal Sample Sizes
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Behrens–Fisher Test
This test...


