Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Cluster Sampling Method01:20

Cluster Sampling Method

Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Weighted Mean00:57

Weighted Mean

While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Test for Homogeneity01:23

Test for Homogeneity

The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can be stated as...
Kruskal-Wallis Test01:19

Kruskal-Wallis Test

The Kruskal-Wallis test, also known as the Kruskal-Wallis H test, serves as a nonparametric alternative to the one-way ANOVA, offering a solution for analyzing the differences across three or more independent groups based on a single, ordinal-dependent variable. This statistical test is particularly valuable in scenarios where the data does not meet the normal distribution assumption required by its parametric counterparts. Kruskal-Wallis test is designed typically to handle ordinal data or...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Wald-Wolfowitz Runs Test II01:17

Wald-Wolfowitz Runs Test II

The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Physical exercise and subjective wellbeing among Chinese students in South Korea: the mediating roles of exercise-related psychological needs satisfaction and psychological resilience.

Frontiers in psychology·2026
Same author

Conformity to stochastic invariance of microstructure influences mechanical competence of trabecular bone.

Bone·2026
Same author

A Modified Normalized Power Prior Approach for Bayesian Adaptive Borrowing in Item Response Theory Models.

Statistics in medicine·2026
Same author

Jieduquyuziyin prescription regulates the differentiation of Tfh cells in systemic lupus erythematosus through lysophosphatidyl ethanolamine metabolism.

Journal of ethnopharmacology·2025
Same author

Common pathological mechanisms and therapeutic strategies in primary Sjogren's syndrome and primary biliary cholangitis: from tissue immune microenvironment to targeted therapy.

Frontiers in immunology·2025
Same author

Modular organization of enhancer network provides transcriptional robustness in mammalian development.

Nucleic acids research·2025

Related Experiment Video

Updated: Jul 10, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

Determining the number of clusters using the weighted gap statistic.

Mingjin Yan1, Keying Ye

  • 1Medtronic Sofamor Danek, 1800 Pyramid Place, Memphis, Tennessee 38132, USA. mingjin.yan@medtronic.com

Biometrics
|April 12, 2007
PubMed
Summary

We introduce new weighted gap methods for estimating the number of clusters in data. These methods, including a multilayer approach, improve accuracy, especially for nested cluster structures.

More Related Videos

A Visual Guide to Sorting Electrophysiological Recordings Using 'SpikeSorter'
10:31

A Visual Guide to Sorting Electrophysiological Recordings Using 'SpikeSorter'

Published on: February 10, 2017

Divergence of Root Microbiota in Different Habitats based on Weighted Correlation Networks
09:49

Divergence of Root Microbiota in Different Habitats based on Weighted Correlation Networks

Published on: September 25, 2021

Related Experiment Videos

Last Updated: Jul 10, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

A Visual Guide to Sorting Electrophysiological Recordings Using 'SpikeSorter'
10:31

A Visual Guide to Sorting Electrophysiological Recordings Using 'SpikeSorter'

Published on: February 10, 2017

Divergence of Root Microbiota in Different Habitats based on Weighted Correlation Networks
09:49

Divergence of Root Microbiota in Different Habitats based on Weighted Correlation Networks

Published on: September 25, 2021

Area of Science:

  • Statistics
  • Data Mining
  • Machine Learning

Background:

  • Estimating the optimal number of clusters is fundamental in cluster analysis.
  • Existing methods like the gap method have limitations, particularly with complex data structures.

Purpose of the Study:

  • To propose novel weighted gap methods for robust cluster number estimation.
  • To introduce a "multilayer" clustering approach for enhanced accuracy.
  • To provide methods applicable to continuous data and any clustering algorithm.

Main Methods:

  • Weighted gap method utilizing weighted within-clusters sum of errors.
  • Difference of difference-weighted (DD-weighted) gap method.
  • Multilayer clustering approach for nested data structures.

Main Results:

  • Proposed weighted gap methods demonstrate improved accuracy over the original gap method.
  • The multilayer approach is particularly effective for detecting nested cluster patterns.
  • Methods validated through simulation studies and real-world data analysis.

Conclusions:

  • Weighted gap and DD-weighted gap methods offer advancements in cluster number estimation.
  • The multilayer approach enhances the detection of intricate data structures.
  • These methods provide flexible and accurate tools for cluster analysis with continuous data.