Related Experiment Video
Updated: Jun 27, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Selective Inference for Hierarchical Clustering
Lucy L Gao1, Jacob Bien2, Daniela Witten3
1Department of Statistics, University of British Columbia.
When groups are identified by clustering, traditional statistical tests inflate the type I error rate. This study introduces a selective inference method to accurately test for mean differences between clusters, controlling for data-driven hypothesis selection.
Area of Science:
- Statistical methodology
- Bioinformatics
- Data science
Background:
- Classical statistical tests assume groups are defined a priori.
- Clustering-based group definitions lead to inflated Type I error rates with traditional tests.
- This issue persists even with independent datasets for clustering and testing.
Purpose of the Study:
- To develop a selective inference approach for testing mean differences between two clusters.
- To control the Type I error rate when hypotheses are data-driven.
- To provide an efficient method for computing exact p-values for hierarchical clustering.
Main Methods:
- Proposed a selective inference procedure to address inflated Type I errors.
- Developed efficient computation of exact p-values for agglomerative hierarchical clustering.
- Validated the method using simulated and single-cell RNA-sequencing data.
Main Results:
- The proposed selective inference method effectively controls the Type I error rate.
- Demonstrated accurate p-value computation for data-driven cluster comparisons.
- Successfully applied the method to real-world single-cell RNA-sequencing data.
Conclusions:
- Selective inference is crucial for valid hypothesis testing after data-driven clustering.
- The developed method offers a statistically sound approach for comparing means between clusters.
- This work has implications for fields utilizing clustering, such as single-cell genomics.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Survival Tree
Building a Survival Tree
Constructing a...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Evolutionary Relationships through Genome Comparisons
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...

