Integrative analysis and variable selection with multiple high-dimensional data sets
Shuangge Ma1, Jian Huang, Xiao Song
1School of Public Health, Yale University, 60 College Street, New Haven, CT 06520, USA. shuangge.ma@yale.edu.
Abstract:
In high-throughput -omics studies, markers identified from analysis of single data sets often suffer from a lack of reproducibility because of sample limitation. A cost-effective remedy is to pool data from multiple comparable studies and conduct integrative analysis. Integrative analysis of multiple -omics data sets is challenging because of the high dimensionality of data and heterogeneity among studies. In this article, for marker selection in integrative analysis of data from multiple heterogeneous studies, we propose a 2-norm group bridge penalization approach. This approach can effectively identify markers with consistent effects across multiple studies and accommodate the heterogeneity among studies. We propose an efficient computational algorithm and establish the asymptotic consistency property. Simulations and applications in cancer profiling studies show satisfactory performance of the proposed approach.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Friedman Two-way Analysis of Variance by Ranks
Gaussian Elimination: Problem Solving
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...

