Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Significance01:37

Statistical Significance

24.0K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
24.0K
Cluster Sampling Method01:20

Cluster Sampling Method

15.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.5K
Identifying Statistically Significant Differences: The F-Test01:14

Identifying Statistically Significant Differences: The F-Test

4.2K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
4.2K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

4.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.6K
Critical Region, Critical Values and Significance Level01:16

Critical Region, Critical Values and Significance Level

13.8K
The critical region, critical value, and significance level are interdependent concepts crucial in hypothesis testing.
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in  probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the...
13.8K
Significance Testing: Overview01:04

Significance Testing: Overview

13.0K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
13.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

High-dimensional test for one-sided hypotheses.

Biostatistics (Oxford, England)·2026
Same author

Targeting the IRE1α/JNK-autophagy axis in microglia attenuates retinal neovascularization in oxygen-induced retinopathy.

Free radical biology & medicine·2026
Same author

Elastic Shape Analysis of Movement Data.

Journal of the American Statistical Association·2026
Same author

Targeting ER Stress of GLP-1 Receptor Agonist in Diabetic Retinopathy.

Biochemical pharmacology·2025
Same author

Determining optimal diet/exercise treatment assignment for patients with symptomatic knee osteoarthritis using baseline gait forces.

Osteoarthritis and cartilage open·2025
Same author

Prognostic significance of CD8+ T cell Spatial Biomarkers in ER+ and ER- breast cancer: A retrospective cohort study.

PLoS medicine·2025

Related Experiment Video

Updated: Mar 27, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.4K

Statistical Significance of Clustering using Soft Thresholding.

Hanwen Huang1, Yufeng Liu2, Ming Yuan3

  • 1Department of Epidemiology and Biostatistics, University of Georgia, Athens, GA 30605.

Journal of Computational and Graphical Statistics : a Joint Publication of American Statistical Association, Institute of Mathematical Statistics, Interface Foundation of North America
|January 13, 2016
PubMed
Summary

A new method improves Statistical Significance of Clustering (SigClust) for high-dimensional data by refining eigenvalue estimation. This enhances cluster detection accuracy, reducing false positives in bioinformatics and cancer genomics.

Keywords:
ClusteringCovariance EstimationHigh DimensionInvariance PrinciplesUnsupervised Learning

More Related Videos

Determination of Aggregate Surface Morphology at the Interfacial Transition Zone ITZ
08:59

Determination of Aggregate Surface Morphology at the Interfacial Transition Zone ITZ

Published on: December 16, 2019

8.8K
Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures
07:05

Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures

Published on: February 15, 2022

3.0K

Related Experiment Videos

Last Updated: Mar 27, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.4K
Determination of Aggregate Surface Morphology at the Interfacial Transition Zone ITZ
08:59

Determination of Aggregate Surface Morphology at the Interfacial Transition Zone ITZ

Published on: December 16, 2019

8.8K
Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures
07:05

Area-based Image Analysis Algorithm for Quantification of Macrophage-fibroblast Cocultures

Published on: February 15, 2022

3.0K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Statistical Learning

Background:

  • Clustering is vital for discoveries but distinguishing true clusters from artifacts is challenging, especially in high-dimensional, low-sample-size data.
  • Existing methods struggle with high dimensionality, and Statistical Significance of Clustering (SigClust) has limitations in eigenvalue estimation.

Purpose of the Study:

  • To address the severe type-I error inflation in SigClust caused by inaccurate eigenvalue estimation in specific high-dimensional scenarios.
  • To develop an improved SigClust method with enhanced accuracy and reliability for cluster evaluation.

Main Methods:

  • Introduced a novel likelihood-based soft thresholding approach for estimating eigenvalues of the covariance matrix under the null multivariate Gaussian distribution.
  • Developed a new metric, the Theoretical Cluster Index, for mathematical analysis of performance.
  • Conducted extensive simulation studies and applied the improved method to cancer genomic data.

Main Results:

  • The novel eigenvalue estimation significantly reduces type-I error inflation compared to the original SigClust method, particularly when few large eigenvalues are present.
  • Mathematical analysis and simulations demonstrate substantial performance improvements in SigClust.
  • The enhanced SigClust method shows practical utility in analyzing cancer genomic datasets.

Conclusions:

  • The proposed soft thresholding method provides a more robust and accurate eigenvalue estimation for SigClust.
  • This advancement leads to a more reliable cluster evaluation tool for high-dimensional, low-sample-size data.
  • The improved SigClust has significant implications for discovery in bioinformatics and related fields.