Related Experiment Video
Updated: Mar 23, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Cross-Clustering: A Partial Clustering Algorithm with Automatic Estimation of the Number of Clusters
Paola Tellaroli1, Marco Bazzi1, Michele Donato2
1Department of Statistical Sciences, University of Padova, Padova, Italy.
This study introduces Cross-clustering (CC), a novel partial clustering algorithm that effectively handles outliers and determines the optimal number of clusters without prior estimation. CC outperforms existing methods in identifying true cluster memberships across diverse datasets.
Area of Science:
- Data Science
- Computational Biology
- Statistics
Background:
- Clustering methods often struggle with outlier detection, require pre-defined cluster numbers, and are sensitive to initialization.
- Existing algorithms lack robustness in identifying inappropriate data partitioning and can yield dependent results.
Purpose of the Study:
- To introduce Cross-clustering (CC), a novel partial clustering algorithm designed to overcome common limitations in data clustering.
- To enhance the accuracy of cluster identification, outlier detection, and membership determination.
Main Methods:
- CC combines principles from Ward's minimum variance and Complete-linkage hierarchical clustering.
- The algorithm was validated against existing methods using both simulated and real-world datasets.
Main Results:
- CC demonstrates superior performance in identifying the correct number of clusters and accurately detecting outliers.
- The method shows improved determination of true cluster memberships compared to traditional algorithms.
- CC proved effective in identifying disease subtypes and gene expression patterns.
Conclusions:
- Cross-clustering (CC) offers a robust and versatile solution for various data clustering challenges, including biological and non-biological applications.
- The CC algorithm is freely available in R, promoting wider adoption and research.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...

