Related Experiment Video
Updated: Jan 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Leveraging external information by guided adaptive shrinkage to improve variable selection in high-dimensional
Mark A van de Wiel1, Wessel N van Wieringen1,2
1Department of Epidemiology and Data Science, Amsterdam Public Health Research Institute, Amsterdam University Medical Centers, Amsterdam, The Netherlands.
Guided adaptive shrinkage methods leverage external co-data to enhance variable selection in high-dimensional, low-sample-size settings. This approach improves prediction accuracy by adapting shrinkage parameters using complementary information, particularly in genomics.
Area of Science:
- Statistics
- Bioinformatics
- Machine Learning
Background:
- High-dimensional data with low sample sizes presents significant variable selection challenges.
- External information, termed 'co-data', can improve variable selection accuracy.
- Co-data, such as variable groupings or prior p-values, is abundant in genomics.
Purpose of the Study:
- To review guided adaptive shrinkage methods that utilize co-data for improved variable selection.
- To discuss the technical aspects and applicability of co-data in prediction models.
- To compare guided shrinkage with other methods like sparse group-lasso.
Main Methods:
- Review of guided adaptive shrinkage methods.
- Adaptation of shrinkage parameters using co-data.
- Comparison with sparse group-lasso for variable selection.
- Integration of co-data learners and spike-and-slab priors for 'do-it-yourself' implementation.
Main Results:
- Guided adaptive shrinkage methods effectively use co-data to enhance variable selection.
- The methodology demonstrates versatility in integrating different co-data types.
- Demonstration of improved variable selection in genetics studies through DIY implementation.
Conclusions:
- Guided adaptive shrinkage offers a powerful framework for variable selection with co-data.
- The methods are applicable across various domains, especially genomics.
- Practical implementation guidance is provided for researchers.
Related Concept Videos
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Survival Tree
Building a Survival Tree
Constructing a...
Variation
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
Outliers and Influential Points
