Related Experiment Video
Updated: Sep 11, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
A methodological study of clustering evaluation in medical data: Exploring internal and stability-based criteria
Azam Orooji1, Farzaneh Kermani2,3, Seyed Mohsen Hosseini4
1Department of Medical Biotechnology, School of Medicine, North Khorasan University of Medical Science (NKUMS), Bojnourd, Iran.
Background:
With the increasing volume and complexity of medical data, clustering methods have become valuable tools for discovering hidden patterns in unlabeled datasets.
Objective:
This study aims to compare the performance of different clustering algorithms for analyzing medical datasets, and evaluate the effectiveness of different clustering quality criteria, including internal, stability and external metrics.
Methods:
A descriptive analysis was conducted on five publicly available medical datasets: Indian Liver Patient Disease (ILPD), Breast Cancer, Pima Indians Diabetes, Statlog (Heart) and Polycystic Ovary Syndrome (PCOS). After preprocessing (handling missing values, normalization, and variable encoding), eight clustering algorithms-Hierarchical Clustering (HC), k-means clustering, Fuzzy ANalYsis (FANNY), Self-Organizing Tree Algorithm (SOTA), Divisive Analysis (DIANA), Partitioning Around Medoids (PAM), Clustering Large Applications (CLARA) and AGglomerative NESting (AGNES) were applied. Clustering quality was evaluated using internal measures (Connectivity, Silhouette Width, Dunn Index), stability measures (Average Proportion of Non-overlap (APN), Average Distance (AD), Average Distance between Means (ADM) and Figure of Merit (FOM)) and external measures (Accuracy and Adjusted Rand Index (ARI)).
Results:
HC and AGNES consistently achieved the best internal and stability performance on the ILPD, Pima, and Statlog datasets, while k-means performed best on the Breast Cancer dataset. On the PCOS dataset, HC, AGNES, k-means, and DIANA showed comparable performance. However, external validation revealed that higher internal and stability scores did not necessarily correspond to greater agreement with clinical class labels, and the optimal algorithm varied across datasets.
Conclusion:
No clustering algorithm consistently outperformed the others across all medical datasets. Although hierarchical methods showed strong internal and stability performance, these metrics alone were insufficient to predict agreement with clinical labels. Therefore, external validation should complement internal and stability metrics to provide a more comprehensive evaluation of clustering performance in medical data analysis.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Comparing the Survival Analysis of Two or More Groups
Data Validation
Key parameters for method validation include:
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...