Related Experiment Video
Updated: Jan 11, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Efficient Hybrid Hierarchical Clustering with Incremental Silhouette Score for Large, Noisy Datasets
Petros Barmpas1, Panagiotis Anagnostou1, Sotiris Tasoulis1
1Department of Computer Science and Biomedical Informatics, University of Thessaly, Papasiopoulou, Lamia 35131, Greece.
None:
This paper introduces a comprehensive framework for clustering analysis, centered on a novel incremental silhouette score calculation designed specifically for hierarchical clustering. This innovative method significantly reduces the computational complexity of silhouette evaluation, transforming the process from O(K N) to effectively O(N) for K hierarchical configurations (demonstrated by an over 100-fold speedup in our tests) making it feasible for large-scale datasets and enabling efficient cluster number estimation within hierarchical clustering scenarios. Building on this, we revisit and enhance the Principal Direction Divisive Partitioning (IPDDP) algorithm, proposing principal component analysis-maximum margin divisive clustering (PCA-MMDC), which utilizes multiple principal components for more accurate data partitioning, and PCA-MMDC-sc, which incorporates a scatter-based cluster selection for improved balance. These are integrated into a hybrid clustering strategy that combines the strengths of incremental silhouette calculation and the enhanced algorithms, allowing for robust cluster identification and effective management of noise and outliers. Experimental results on synthetic and real-world datasets demonstrate notable improvements in clustering accuracy (achieving an average Adjusted Rand Index (ARI) increase of over 10 percentage points on custom noisy synthetic datasets compared to K-Means) and computational efficiency. While the choice of principal components in PCA-MMDC presents a parameter, the overall framework offers a scalable and robust solution for complex clustering tasks, with future work aimed at adaptive parameter selection and extending incremental calculations to other validation metrics.