Estimating the number of clusters via a corrected clustering instability
Jonas M B Haslbeck1, Dirk U Wulff2,3
1Psychological Methods Group, University of Amsterdam, Amsterdam, The Netherlands.
Summary
We developed a new method to accurately select the number of clusters (k) in cluster analysis. Our corrected instability measure improves upon existing techniques, especially for larger datasets.
Area of Science:
- Cluster analysis
- Data mining
- Statistical modeling
Background:
- Determining the optimal number of clusters (k) is crucial for accurate cluster analysis.
- Existing instability-based methods can be influenced by cluster size distributions, limiting their effectiveness, particularly for large k.
- A need exists for robust methods to select k that are less sensitive to cluster size variations.
Purpose of the Study:
- To improve instability-based methods for selecting the number of clusters (k) in cluster analysis.
- To introduce a corrected clustering distance that mitigates the impact of cluster size distribution on instability.
- To compare the performance of model-based and model-free approaches for determining cluster instability.
Main Methods:
- Development of a corrected clustering distance metric.
- Evaluation of the corrected instability measure against existing methods across a range of k values.
- Comparative analysis of model-based and model-free cluster instability determination.
Main Results:
- The corrected instability measure demonstrates superior performance compared to current instability-based methods for all tested values of k.
- The proposed method effectively overcomes limitations of existing techniques, particularly for large k.
- Model-based and model-free approaches for determining cluster instability exhibit comparable performance.
Conclusions:
- The corrected clustering distance offers a more reliable approach for selecting the optimal number of clusters (k).
- This advancement enhances the robustness of instability-based methods in cluster analysis.
- The developed method is available as the R-package cstab for practical application.
Related Concept Videos
Cluster Sampling Method
13.8K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.8K
Estimating Population Mean with Unknown Standard Deviation
8.6K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.6K
Estimating Population Standard Deviation
3.2K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.2K
Estimating Population Mean with Known Standard Deviation
9.4K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
9.4K
Chi-square Analysis
42.2K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
42.2K
Distributions to Estimate Population Parameter
4.9K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.9K


