Related Experiment Video
Updated: Jan 21, 2026

JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Gaussian mixture copulas for high-dimensional clustering and dependency-based subtyping
Siva Rajesh Kasa1, Sakyajit Bhattacharya2, Vaibhav Rajan1
1Department of Information Systems and Analytics, School of Computing, National University of Singapore, 117418 Singapore.
Motivation:
The identification of sub-populations of patients with similar characteristics, called patient subtyping, is important for realizing the goals of precision medicine. Accurate subtyping is crucial for tailoring therapeutic strategies that can potentially lead to reduced mortality and morbidity. Model-based clustering, such as Gaussian mixture models, provides a principled and interpretable methodology that is widely used to identify subtypes. However, they impose identical marginal distributions on each variable; such assumptions restrict their modeling flexibility and deteriorates clustering performance.
Results:
In this paper, we use the statistical framework of copulas to decouple the modeling of marginals from the dependencies between them. Current copula-based methods cannot scale to high dimensions due to challenges in parameter inference. We develop HD-GMCM, that addresses these challenges and, to our knowledge, is the first copula-based clustering method that can fit high-dimensional data. Our experiments on real high-dimensional gene-expression and clinical datasets show that HD-GMCM outperforms state-of-the-art model-based clustering methods, by virtue of modeling non-Gaussian data and being robust to outliers through the use of Gaussian mixture copulas. We present a case study on lung cancer data from TCGA. Clusters obtained from HD-GMCM can be interpreted based on the dependencies they model, that offers a new way of characterizing subtypes. Empirically, such modeling not only uncovers latent structure that leads to better clustering but also meaningful clinical subtypes in terms of survival rates of patients.
Availability And Implementation:
An implementation of HD-GMCM in R is available at: https://bitbucket.org/cdal/hdgmcm/.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
More Related Videos
Related Concept Videos
Mixtures of Acids
A Mixture of a Strong Acid and a Weak Acid
In a mixture of a strong acid and a weak acid, the strong acid dissociates completely and becomes a source of almost all the hydronium ions...
Mixtures of Acids
In a strong and weak acid mixture, the strong acid dissociates completely and becomes a source of almost all the hydronium ions present in the solution. In contrast, the weak acid shows...
Racemic Mixtures and the Resolution of Enantiomers
Dimensional Analysis
Conversion Factors and Dimensional Analysis
The unit...
Adrenergic Receptors: ɑ Subtype
Adrenaline ≥ Noradrenaline >> Isoprenaline
α-adrenoceptors are further divided into α1 and α2-adrenoceptors.
α1-Adrenoceptors: These receptors are located postsynaptically on the effector organs and cause constriction of smooth muscle mediated by activation of phospholipase...
Adrenergic Receptors: β Subtype
Isoprenaline > Adrenaline > Noradrenaline
Neurotransmitter binding to these receptors causes activation of adenylyl cyclase resulting in increased concentrations of cAMP and modulation of calcium ion channels within the cell. They are further classified into β1, β2, and β3 subtypes.
β1-adrenoceptors: β1-adrenoceptors...

