Related Experiment Video
Updated: Aug 20, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
A comparison of central-tendency and interconnectivity approaches to clustering multivariate data with irregular
Mark Tozer1,2, David Keith2
1NSW Department of Environment Parramatta New South Wales Australia.
Questions:
Most clustering methods assume data are structured as discrete hyperspheroidal clusters to be evaluated by measures of central tendency. If vegetation data do not conform to this model, then vegetation data may be clustered incorrectly. What are the implications for cluster stability and evaluation if clusters are of irregular shape or density?
Location:
Southeast Australia.
Methods:
We define misplacement as the placement of a sample in a cluster other than (distinct from) its nearest neighbor and hypothesize that optimizing homogeneity incurs the cost of higher rates of misplacement. Chameleon is a graph-theoretic algorithm that emphasizes interconnectivity and thus is sensitive to the shape and distribution of clusters. We contrasted its solutions with those of traditional nonhierarchical and hierarchical (agglomerative and divisive) approaches.
Results:
Chameleon-derived solutions had lower rates of misplacement and only marginally higher heterogeneity than those of k-means in the range of 15-60 clusters, but their metrics converged with larger numbers of clusters. Solutions derived by agglomerative clustering had the best metrics (and divisive clustering the worst) but both produced inferior high-level solutions to those of Chameleon by merging distantly-related clusters.
Conclusions:
Graph-theoretic algorithms, such as Chameleon, have an advantage over traditional algorithms when data exhibit discontinuities and variable structure, typically producing more stable solutions (due to lower rates of misplacement) but scoring lower on traditional metrics of central tendency. Advantages are less obvious in the partitioning of data from continuous gradients; however, graph-based partitioning protocols facilitate the hierarchical integration of solutions.
Related Concept Videos
What is Central Tendency?
The central tendency is the most conventionally used data characteristic. It is a...
Midrange
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Central Tendency: Analysis
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Skewness
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...

