Related Experiment Video
Updated: Aug 12, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
DDCAL: Evenly Distributing Data into Low Variance Clusters Based on Iterative Feature Scaling
Marian Lux1,2, Stefanie Rinderle-Ma3
1Research Group Workflow Systems and Technology, University of Vienna, Vienna, Austria.
Abstract:
This work studies the problem of clustering one-dimensional data points such that they are evenly distributed over a given number of low variance clusters. One application is the visualization of data on choropleth maps or on business process models, but without over-emphasizing outliers. This enables the detection and differentiation of smaller clusters. The problem is tackled based on a heuristic algorithm called DDCAL (1d distribution cluster algorithm) that is based on iterative feature scaling which generates stable results of clusters. The effectiveness of the DDCAL algorithm is shown based on 5 artificial data sets with different distributions and 4 real-world data sets reflecting different use cases. Moreover, the results from DDCAL, by using these data sets, are compared to 11 existing clustering algorithms. The application of the DDCAL algorithm is illustrated through the visualization of pandemic and population data on choropleth maps as well as process mining results on process models.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Scaling
Coefficient of Variation
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...

