Related Experiment Video
Updated: May 15, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
VICatMix: variational Bayesian clustering and variable selection for discrete biomedical data
Jackie Rao1, Paul D W Kirk1,2,3
1MRC Biostatistics Unit, University of Cambridge, Cambridge, CB2 0SR, United Kingdom.
Summary:
Effective clustering of biomedical data is crucial in precision medicine, enabling accurate stratification of patients or samples. However, the growth in availability of high-dimensional categorical data, including 'omics data, necessitates computationally efficient clustering algorithms. We present VICatMix, a variational Bayesian finite mixture model designed for the clustering of categorical data. The use of variational inference (VI) in its training allows the model to outperform competitors in terms of computational time and scalability, while maintaining high accuracy. VICatMix furthermore performs variable selection, enhancing its performance on high-dimensional, noisy data. The proposed model incorporates summarization and model averaging to mitigate poor local optima in VI, allowing for improved estimation of the true number of clusters simultaneously with feature saliency. We demonstrate the performance of VICatMix with both simulated and real-world data, including applications to datasets from The Cancer Genome Atlas, showing its use in cancer subtyping and driver gene discovery. We demonstrate VICatMix's potential utility in integrative cluster analysis with different 'omics datasets, enabling the discovery of novel disease subtypes.
Availability And Implementation:
VICatMix is freely available as an R package via CRAN, incorporating C++ for faster computation, at https://CRAN.R-project.org/package=VICatMix.
Related Concept Videos
Biostatistics: Overview
Discrete variables are...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Statistical Methods for Analyzing Epidemiological Data
Statistical Software for Data Analysis and Clinical Trials

