Method for Determining the Optimal Number of Clusters Based on Agglomerative Hierarchical Clustering
IEEE Transactions on Neural Networks and Learning Systems
|January 24, 2017
Summary
Determining the optimal number of clusters is key for cluster analysis quality. A new validity index and method using sample geometry concepts improve cluster number determination for various data structures.
Area of Science:
- Data Science
- Machine Learning
- Statistics
Background:
- Accurate cluster number determination is vital for effective cluster analysis.
- Existing methods may not adequately capture geometric properties of sample clusters.
- Evaluating clustering quality requires robust validity indices.
Purpose of the Study:
- To introduce a novel clustering validity index based on sample geometry.
- To propose a method for determining the optimal number of clusters using agglomerative hierarchical clustering (AHC).
- To evaluate the proposed index and method across diverse dataset structures.
Main Methods:
- Definition of sample clustering dispersion and synthesis degrees.
- Design of a new clustering validity index.
- Implementation of a method for optimal cluster number selection with AHC.
Main Results:
- The proposed index effectively evaluates clustering results from AHC.
- The method successfully determines the optimal number of clusters for various data types (linear, manifold, annular, convex).
- Experimental validation confirms the index and method's performance.
Conclusions:
- The new geometric-based validity index offers a reliable approach to cluster analysis.
- The proposed method enhances the determination of the optimal number of clusters.
- This work provides a valuable tool for improving clustering quality across different data geometries.
Related Concept Videos
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K
Sampling Plans
1.1K
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
1.1K
Kruskal-Wallis Test
1.4K
The Kruskal-Wallis test, also known as the Kruskal-Wallis H test, serves as a nonparametric alternative to the one-way ANOVA, offering a solution for analyzing the differences across three or more independent groups based on a single, ordinal-dependent variable. This statistical test is particularly valuable in scenarios where the data does not meet the normal distribution assumption required by its parametric counterparts. Kruskal-Wallis test is designed typically to handle ordinal data or...
1.4K
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.3K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.3K
One-Way ANOVA: Unequal Sample Sizes
6.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.8K


