Related Experiment Video
Updated: Jan 25, 2026

Stable Aqueous Suspensions of Manganese Ferrite Clusters with Tunable Nanoscale Dimension and Composition
Published on: February 5, 2022
On Perfect Clustering of High Dimension, Low Sample Size Data
Clustering high-dimensional data is challenging. A new measure, MADD, improves clustering performance and cluster number estimation in high dimension, low sample size (HDLSS) situations.
Area of Science:
- Statistics
- Data Mining
- Machine Learning
Background:
- Traditional clustering algorithms struggle with high-dimensional, low-sample-size (HDLSS) data due to distance concentration and neighborhood structure issues.
- Euclidean distance-based methods are particularly affected, leading to poor performance in these scenarios.
Purpose of the Study:
- To introduce and evaluate a novel data-driven dissimilarity measure, MADD (Measure of Aggregated Dissimilarity Distance), designed to overcome HDLSS challenges in clustering.
- To demonstrate the effectiveness of MADD in improving clustering performance and cluster number estimation for high-dimensional datasets.
Main Methods:
- Development of the MADD dissimilarity measure, leveraging the distance concentration phenomenon in high dimensions.
- Theoretical and numerical studies to validate MADD's performance against traditional distance measures.
- Adaptation and creation of cluster number estimation algorithms, incorporating MADD for enhanced high-dimensional performance.
- Analysis of simulated and real-world datasets to showcase MADD's practical utility.
Main Results:
- Clustering algorithms utilizing MADD exhibit superior performance in high-dimensional, low-sample-size (HDLSS) settings compared to those using standard distance functions.
- MADD effectively addresses the adverse effects of distance concentration and neighborhood structure violations.
- Existing cluster number estimation algorithms show improved accuracy when integrated with MADD.
- A new, consistent estimator for the number of clusters in the HDLSS regime was developed and validated.
Conclusions:
- The MADD dissimilarity measure offers a robust solution for clustering high-dimensional data, particularly in HDLSS scenarios.
- MADD enhances the performance of both clustering algorithms and cluster number estimation techniques.
- This approach provides a valuable tool for effective cluster analysis in complex, high-dimensional datasets.
More Related Videos
08:21A Simple Method for the Size Controlled Synthesis of Stable Oligomeric Clusters of Gold Nanoparticles under Ambient Conditions
Published on: February 5, 2016
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
Catalytically Perfect Enzymes
Most enzymes...
Support Reactions in Three Dimensions
Ball and Socket Joint is one of the supports allowing free rotation about any axis. This freedom of rotation is...