Related Experiment Video
Updated: Nov 30, 2025

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.1K
Weighted dimensionality reduction and robust Gaussian mixture model based cancer patient subtyping from gene
1Machine Learning Lab, Department of Electronics and Communication Engineering, National Institute of Technology, Srinagar, JK, India.
Journal of Biomedical Informatics
|November 14, 2020
Summary
This study introduces a robust clustering pipeline to improve cancer patient subtyping from gene expression data. The method enhances subtype separation, leading to better statistical and clinical significance for treatment targets.
Area of Science:
- Computational biology
- Cancer genomics
- Biostatistics
Background:
- Cancer's heterogeneity requires patient subtyping for effective treatment.
- Gene expression data presents computational challenges due to high dimensionality, noise, and outliers.
- Existing subtyping methods often yield overlapping survival plots, hindering clear distinction between cancer subtypes.
Purpose of the Study:
- To develop a robust computational framework for improved cancer patient subtyping.
- To enhance the separation and statistical significance of discovered cancer subtypes.
- To identify potential therapeutic targets through pathway analysis.
Main Methods:
- A robust clustering pipeline integrating dimensionality reduction and Gaussian mixture model-based clustering.
- Weighted gene expression matrix using median absolute deviation (MAD) before t-distributed Stochastic Neighbor Embedding (t-SNE) for dimensionality reduction.
- Quantification of subtype separation using minimum pairwise Kaplan-Meier (KM) plot separation, minimum hazard ratio, and a novel cumulative survival separation metric.
Main Results:
- The proposed method significantly improved subtype separation compared to Consensus Clustering, CNMF, Spectrum, and NEMO across five TCGA datasets.
- For Glioblastoma Multiforme (GBM), the method achieved a 61-day survival difference between subtypes, outperforming existing methods.
- Utilizing MAD scores notably enhanced subtype separation, and hazard ratio analysis confirmed superior performance.
Conclusions:
- The integration of median absolute deviation (MAD) and a robust clustering methodology effectively improves cancer subtype separation.
- Enhanced subtype separation offers greater statistical and clinical significance for understanding cancer heterogeneity.
- The findings provide a foundation for identifying more precise therapeutic targets in cancer treatment.

