Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Significance01:50

Statistical Significance

21.1K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
21.1K
Cluster Sampling Method01:20

Cluster Sampling Method

14.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.0K
Introduction to Test of Independence01:21

Introduction to Test of Independence

2.9K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.9K
Test for Homogeneity01:23

Test for Homogeneity

2.4K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.4K
Significance Testing: Overview01:04

Significance Testing: Overview

11.5K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
11.5K
Determination of Expected Frequency01:08

Determination of Expected Frequency

2.5K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

CD38⁺ endothelial remodeling marks spatially patterned vasculopathy in rapidly advancing periodontitis and peri-implantitis.

Nature communications·2026
Same author

GARD: Genomic Data-Based Drug Repurposing in Head and Neck Cancer with Large Language Model Validation.

Cancers·2026
Same author

GARD: Genomic Data based Drug Repurposing in Head and Neck Cancer with Large Language Model Validation.

bioRxiv : the preprint server for biology·2026
Same author

Differential roles of Rad18 in repressing carcinogen- and oncogene-driven mutagenesis <i>in vivo</i>.

NAR cancer·2026
Same author

Decoding longitudinal microbiome trajectories: an interpretable machine learning approach for biomarker discovery and prediction.

Briefings in bioinformatics·2025
Same author

CD38<sup>+</sup> Endothelial Remodeling Defines Spatially Diverse Vasculopathy Programs in Rapidly Advancing Oral Inflammation.

bioRxiv : the preprint server for biology·2025

Related Experiment Video

Updated: Jan 17, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.3K

Statistical significance of clustering for count data.

Yifan Dai1, Di Wu1,2, Yufeng Liu3

  • 1Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA.

Biometrics
|September 19, 2025
PubMed
Summary

We introduce SigClust-DEV, a novel method for assessing cluster significance in count data, outperforming existing approaches. This tool enhances subgroup identification in genomics and healthcare data by addressing statistical uncertainty.

Keywords:
clustering significancedimension reductionelectronic health recordsgeneralized principal component analysissingle-cell RNA-sequencing data

More Related Videos

Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
13:55

Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization

Published on: February 3, 2013

18.9K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

493

Related Experiment Videos

Last Updated: Jan 17, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.3K
Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
13:55

Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization

Published on: February 3, 2013

18.9K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

493

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Statistical Genetics

Background:

  • Clustering is vital for identifying subgroups in biomedical research, but existing methods often overlook statistical uncertainty, leading to spurious clusters.
  • The Statistical Significance of Clustering (SigClust) method assesses cluster significance in high-dimensional data but is limited to continuous data and can lack power for non-Gaussian distributions.

Purpose of the Study:

  • To develop a novel method, SigClust-DEV, for evaluating the statistical significance of clusters specifically in discrete count data.
  • To address the limitations of existing SigClust methods in handling count data and non-Gaussian distributions.

Main Methods:

  • SigClust-DEV was developed to assess cluster significance in high-dimensional count data.
  • Extensive simulations were conducted to compare SigClust-DEV against existing SigClust variations across diverse count distributions.
  • The method was applied to single-cell RNA sequencing (scRNA) data and electronic health records (EHRs) for real-world validation.

Main Results:

  • SigClust-DEV demonstrated superior performance compared to existing SigClust approaches in simulations across various count distributions.
  • The method successfully identified meaningful latent cell types in Hydra scRNA data.
  • SigClust-DEV effectively identified significant patient subgroups within cancer EHR data.

Conclusions:

  • SigClust-DEV is a powerful and statistically robust method for assessing cluster significance in count data.
  • The method enhances subgroup discovery in complex biomedical datasets like scRNA and EHR data.
  • SigClust-DEV overcomes limitations of previous methods, offering improved statistical power and accuracy for discrete data analysis.