Related Experiment Videos
Statistical significance for hierarchical clustering in genetic association and microarray expression studies
Mark A Levenstien1, Yaning Yang, Jürg Ott
1Laboratory of Statistical Genetics, Rockefeller University, New York, NY 10021, United States. markl@linkage.rockefeller.edu
BMC Bioinformatics
|December 12, 2003
Summary
Hierarchical clustering in genetics can overstate significance. Focusing on the best single result ignores the clustering process, potentially leading to false positives in genetic association and gene expression studies.
Area of Science:
- Genetics
- Bioinformatics
- Statistical Analysis
Background:
- Vast datasets in molecular genetics (gene expression, haplotypes) necessitate data reduction for interpretation.
- Hierarchical clustering is commonly used to group similar observations into fewer classes.
- Researchers often select the most significant result at a specific clustering step, overlooking the process itself.
Purpose of the Study:
- To develop a method for assessing the overall statistical significance of hierarchically clustered data.
- To address the issue of inflated significance when focusing on a single optimal clustering outcome.
Main Methods:
- Propose using the strongest result (smallest p-value) as an experiment-wise statistic.
- Evaluate the significance level of this statistic for a global assessment.
- Applied the approach to haplotype association and microarray expression datasets.
Main Results:
- Selecting a single clustering step often yields significance levels that are too small.
- The proposed method provides a more accurate global assessment of statistical significance.
- Over-reliance on one clustering outcome can lead to formally significant results for non-significant overall experiments.
Conclusions:
- The clustering process itself impacts statistical significance and must be incorporated.
- A single, optimal clustering result may not reflect true experimental significance.
- Accurate interpretation of genetic data requires methods that account for the entire clustering procedure.