Related Experiment Video
Updated: Sep 17, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Benchmarking validity indices for evolutionary K-means clustering performance
Abiodun M Ikotun1, Faustin Habyarimana1, Absalom E Ezugwu2
1School of Mathematics, Statistics and Computer Science, University of KwaZulu- Natal, KwaZulu-Natal, Pietermaritzburg Campus, Durban, South Africa.
This study evaluated internal validity indices for Evolutionary K-Means clustering. The Calinski-Harabasz (CH) and Silhouette indices proved most effective for automatic clustering tasks.
Area of Science:
- Data Science
- Machine Learning
- Artificial Intelligence
Background:
- K-Means clustering requires a predefined number of clusters, limiting its use in automatic data analysis.
- Evolutionary K-Means (E-KM) algorithms integrate metaheuristics to overcome K-Means limitations, using internal validity indices for automatic cluster determination.
- The performance of internal validity indices is data-dependent, impacting the reliability of E-KM outcomes.
Purpose of the Study:
- To evaluate the performance of fifteen internal validity indices within the Enhanced Firefly Algorithm-K-Means (FA-K-Means) framework.
- To identify the most effective internal validity indices for automatic clustering tasks using an evolutionary approach.
- To provide practical guidance on selecting fitness functions for E-KM algorithms.
Main Methods:
- The study employed the Enhanced Firefly Algorithm-K-Means (FA-K-Means) framework, combining Firefly metaheuristics with K-Means.
- Fifteen distinct internal validity indices were assessed as fitness functions within the FA-K-Means framework.
- Performance was evaluated across a variety of real-life and synthetic datasets with diverse structural properties.
Main Results:
- The Calinski-Harabasz (CH) index demonstrated consistently strong performance across various datasets.
- The Silhouette index also showed robust and reliable performance in determining optimal clustering configurations.
- Other evaluated indices exhibited variable effectiveness, often dependent on specific dataset characteristics.
Conclusions:
- The Calinski-Harabasz (CH) and Silhouette indices are recommended for use as fitness functions in Evolutionary K-Means algorithms.
- These indices offer more reliable and consistent clustering performance for automatic clustering tasks.
- The findings offer practical insights for researchers and practitioners selecting validity indices in E-KM applications.
Related Concept Videos
Kendall's Coefficient of Concordance
Expected Frequencies in Goodness-of-Fit Tests
Goodness-of-Fit Test
Quantifying and Rejecting Outliers: The Grubbs Test
Spearman's Rank Correlation Test
Spearman's test calculates...
Kendall's Tau Test
A τ value...

