Testing for a difference in means of a single feature after clustering
1Department of Biomedical Data Science, Stanford University, 450 Serra Mall, Stanford, CA 94305, United States.
Biostatistics (Oxford, England)
|December 31, 2024
Summary
We developed a new statistical test to accurately compare feature means between two clusters from hierarchical or k-means clustering. This method controls Type I errors, enhancing cluster validation for biological data analysis.
Area of Science:
- Computational biology
- Statistical genetics
- Bioinformatics
Background:
- Interpreting and validating clustered observations is crucial for many applications.
- Classical hypothesis tests for comparing feature means between clusters inflate Type I error rates.
- Accurate cluster validation is essential for reliable data interpretation in fields like single-cell genomics.
Purpose of the Study:
- To propose a novel hypothesis test for comparing feature means between two clusters.
- To address the inflated Type I error rates associated with classical tests in cluster analysis.
- To provide an efficient and reliable method for validating clusters derived from hierarchical or k-means algorithms.
Main Methods:
- Development of a new statistical test for the difference in means between two clusters.
- The test is designed for use with hierarchical clustering and k-means clustering outputs.
- Evaluation of the test's performance through simulations and application to real-world single-cell RNA-sequencing data.
Main Results:
- The proposed test effectively controls the selective Type I error rate in finite samples.
- The method is computationally efficient.
- Simulations demonstrate the test's validity and statistical power.
Conclusions:
- The new test offers a statistically sound approach for validating cluster comparisons.
- It provides a reliable alternative to classical methods that suffer from inflated Type I errors.
- The test is applicable to diverse datasets, including single-cell RNA-sequencing, improving biological insights.
Related Concept Videos
Test for Homogeneity
1.9K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
1.9K
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
Significance Testing: Overview
3.3K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
3.3K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K
Bonferroni Test
2.7K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.7K
Behrens–Fisher Test
66
The Behrens-Fisher test is a statistical method designed to address the Behrens-Fisher problem, which arises when comparing the means of two normally distributed populations with unequal variances. Unlike the Student's t-test, which assumes equal variances, the Behrens-Fisher test allows for mean comparison without this restrictive assumption. This flexibility makes it particularly valuable in scenarios where two independent samples exhibit normality but lack variance homogeneity.
This test...
This test...
66


