Related Experiment Video
Updated: Apr 8, 2026

Basics of Multivariate Analysis in Neuroimaging Data
Published on: July 24, 2010
Variational Bayes for High-Dimensional Multi-Source Heterogeneous Data With Sparse Priors
Wenting Liu1, Lu Luo1, Huiqiong Li1
1Yunnan Key Laboratory of Statistical Modeling and Data Analysis, Yunnan University, Kunming City, Yunnan Province, China.
Abstract:
High-dimensional data is becoming increasingly prevalent in scientific fields such as genomics, economics, and medicine. However, recent studies have indicated that such data is often heterogeneous. Most existing research focuses either on high-dimensional multi-source data or on isolated heterogeneous datasets, leaving a significant gap for joint modeling and inference of high-dimensional multi-source heterogeneous data. In this paper, we employ a Bayesian method to address the estimation problem associated with high-dimensional multi-source heterogeneous data, aiming to extract shared features across all subpopulations while also examining the unique heterogeneity within each subpopulation. We introduce a scalable and interpretable Bayesian model for multi-source heterogeneous linear data that employs a sparsity-inducing spike-and-slab prior, featuring a Laplace slab and a Dirac spike. To address the computational challenges associated with the posterior, we implement a mean-field variational approximation that utilizes a factorizable family of spike-and-slab distributions. Our method overcomes the high computational cost of Gibbs sampling while preserving valuable features, providing a posterior distribution for the parameters and offering a natural mechanism for variable selection via posterior inclusion probabilities. Through simulation studies and an application to five cancer datasets in real-world the Cancer Genome Atlas (TCGA), we demonstrate the effectiveness of our variational Bayesian approach. Our results show the advantages of our method in terms of computational efficiency and scalability compared to Gibbs sampling method and penalized frequentist methods for integrative analysis, making it well-suited for analyzing high-dimensional multi-source heterogeneous data. The proposed variational Bayesian algorithms have been implemented in the R package VBMS, which is publicly available on CRAN.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Multi-input and Multi-variable systems
In the absence of...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Distributions to Estimate Population Parameter
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Probability Distributions
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...

