Related Experiment Video
Updated: Nov 17, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Head-to-head comparison of clustering methods for heterogeneous data: a simulation-driven benchmark.
Gregoire Preud'homme1,2, Kevin Duarte1, Kevin Dalleau3
1Centre d'Investigations Cliniques Plurithématique 1433, INSERM 1116, CHRU de Nancy, Université de Lorraine, Nancy, France.
Choosing unsupervised machine learning for mixed data is hard. Model-based methods like Kamila, LCM, and K-prototypes generally outperform others for heterogeneous datasets.
Area of Science:
- Machine Learning
- Data Science
- Bioinformatics
Background:
- Clustering mixed data (continuous and categorical variables) presents challenges.
- Selecting appropriate unsupervised machine learning methods is crucial for accurate analysis.
Purpose of the Study:
- To benchmark clustering strategies for mixed data.
- To compare model-based and distance-based methods using simulated and real-world clinical trial data.
Main Methods:
- Evaluated 4 model-based (Kamila, Latent Class Analysis, Latent Class Model [LCM], Clustering by Mixture Modeling) and 5 distance/dissimilarity-based methods (K-prototypes, Gower distance, Unsupervised Extra Trees dissimilarity with hierarchical clustering or Partitioning Around Medoids).
- Assessed performance using Adjusted Rand Index (ARI) on simulated data across 7 scenarios.
- Applied methods to EPHESUS heart failure clinical trial data.
Main Results:
- K-prototypes, Kamila, and LCM models demonstrated superior performance.
- Model-based methods generally outperformed dissimilarity-based methods (Partitioning Around Medoids, Hierarchical Clustering).
- LCM showed promise in clinical data for profile differentiation, prognosis (C-index), and identifying treatment benefit subgroups.
Conclusions:
- Significant performance differences exist among clustering algorithms for mixed data.
- Model-based methods (Kamila, LCM) and K-prototypes are recommended for heterogeneous data analysis.
- LCM offers valuable insights in clinical settings for patient subgroup identification and prognosis.
More Related Videos
05:12ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Comparing the Survival Analysis of Two or More Groups
Test for Homogeneity
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Causes of Similarity-Dissimilarity Effect
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...