A modified and weighted Gower distance-based clustering analysis for mixed type data: a simulation and empirical
Pinyan Liu1, Han Yuan2, Yilin Ning2
1Centre for Quantitative Medicine, Duke-NUS Medical School, 8 College Road, Singapore, 169857, Singapore. pinyanliu@u.duke.nus.edu.
A new clustering algorithm, DAFI, effectively analyzes mixed-type clinical data by incorporating feature importance. This approach improves clustering accuracy and reveals significant health associations, such as periodontitis and cardiovascular diseases.
Area of Science:
- Data Science
- Biostatistics
- Computational Biology
Background:
- Traditional clustering methods struggle with mixed-type data (continuous and categorical variables).
- Real-world clinical datasets frequently contain a mix of data types, necessitating advanced analytical approaches.
- Existing techniques lack adaptability and interpretability for complex, heterogeneous datasets.
Purpose of the Study:
- To introduce a novel clustering technique, DAFI (Distance-based Algorithm with Feature Importance), designed for mixed-type data.
- To enhance clustering compatibility, adaptability, and interpretability in clinical research.
- To identify distinct patient clusters and their associated health profiles using real-world data.
Main Methods:
- Developed a modified Gower distance metric incorporating feature importance weights.
- Evaluated the DAFI algorithm on simulated datasets and the National Health and Nutrition Examination Survey (NHANES) data.
- Compared DAFI's performance against 13 existing clustering techniques using Adjusted Rand Index (ARI) and silhouette scores.
Main Results:
- DAFI consistently outperformed baseline methods in simulation studies, particularly with redundant features (ARI).
- DAFI achieved the highest silhouette score (0.79) on NHANES data, identifying four distinct health clusters.
- Feature importance highlighted cardiovascular disease (CVD) risk factors in cluster formation.
- Analysis revealed a significant association between periodontitis (PD) and CVDs (adjusted OR 1.95).
Conclusions:
- DAFI demonstrates superior performance over traditional clustering baselines for both simulated and real-world mixed-type data.
- The algorithm effectively captures cluster characteristics by weighting feature importance, crucial for clinical relevance.
- DAFI provides a robust and interpretable solution for mixed-type data clustering in healthcare settings.
More Related Videos
05:12ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
08:49Printed Glycan Array: A Sensitive Technique for the Analysis of the Repertoire of Circulating Anti-carbohydrate Antibodies in Small Animals
Published on: February 14, 2019
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Quantifying and Rejecting Outliers: The Grubbs Test
Comparing the Survival Analysis of Two or More Groups
One-Way ANOVA: Unequal Sample Sizes
Friedman Two-way Analysis of Variance by Ranks
