Related Experiment Video
Updated: Jan 2, 2026

A Streamlined Approach for Mass Spectrometry-Based Proteomics Using Selected Tissue Regions
Published on: April 18, 2025
Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer
Oliver M Crook1,2,3, Laurent Gatto4, Paul D W Kirk3,5
1Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK.
This study introduces an enhanced Dirichlet Process mixture model for clustering that incorporates variable selection and Bayesian model averaging. The new method offers competitive performance and computational advantages for analyzing complex biological data, including cancer transcriptomics and proteomics.
Area of Science:
- Computational statistics
- Bioinformatics
- Machine learning
Background:
- Dirichlet Process (DP) mixture models are widely used for model-based clustering, enabling inference of the number of clusters.
- The sequential updating and greedy search (SUGS) algorithm provides efficient approximate Bayesian inference for DP mixture models.
- Existing methods often lack integrated variable selection and may not fully leverage Bayesian model averaging (BMA).
Purpose of the Study:
- To extend the SUGS algorithm for DP mixture models to incorporate variable selection for clustering.
- To demonstrate the advantages of using Bayesian model averaging (BMA) over Bayesian model selection (BMS) in this context.
- To provide a computationally efficient and robust method for analyzing high-dimensional biological data.
Main Methods:
- Extension of the sequential updating and greedy search (SUGS) algorithm to include variable selection.
- Application of Bayesian model averaging (BMA) for improved inference in DP mixture models.
- Implementation in an open-source R package (sugsvarsel) with C++ acceleration and parallel processing.
Main Results:
- The proposed method demonstrates competitive performance against state-of-the-art approaches in simulation studies and cancer transcriptomics examples.
- Significant computational benefits are observed compared to traditional Markov chain Monte Carlo methods.
- Successful application to The Cancer Genome Atlas (TCGA) reverse-phase protein array (RPPA) data for pan-cancer proteomic characterization.
Conclusions:
- The enhanced SUGS algorithm with variable selection and BMA offers a powerful and efficient tool for model-based clustering.
- The method provides valuable insights into complex biological datasets, such as cancer proteomic profiles.
- The open-source R package sugsvarsel facilitates broader adoption and accelerates research in computational biology and statistics.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Cancer Survival Analysis
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Distributions to Estimate Population Parameter
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

