Related Experiment Video
Updated: Jan 6, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.5K
Adaptive penalization in high-dimensional regression and classification with external covariates using variational
Britta Velten1, Wolfgang Huber1
1Genome Biology Unit, European Molecular Biology Laboratory, Meyerhofstr. 1, 69117 Heidelberg, Germany.
Biostatistics (Oxford, England)
|October 10, 2019
Summary
This study introduces a novel penalized regression method that uses external covariates to adapt feature penalization. The approach improves prediction accuracy and model interpretability by leveraging group-specific information.
Area of Science:
- Statistics
- Machine Learning
- Bioinformatics
Background:
- Penalized regression methods like Lasso and ridge regression are standard for high-dimensional data.
- The relative strength of penalization is often implicitly determined by predictor scales, ignoring valuable external information.
- Existing methods do not fully utilize available covariate information to adapt penalization strategies.
Purpose of the Study:
- To develop a data-driven method for adapting penalized regression by incorporating external covariates.
- To differentially penalize feature groups based on covariates and their information content.
- To enhance prediction performance and model interpretability in high-dimensional settings.
Main Methods:
- A Bayesian approach is employed to combine shrinkage and feature selection.
- The method introduces a scalable optimization scheme for penalized regression.
- Feature groups are defined by external covariates, allowing for adaptive penalization strengths.
Main Results:
- Simulations show accurate recovery of true effect sizes and sparsity patterns within feature groups.
- The proposed method demonstrates improved prediction performance when groups have varying dynamic ranges.
- Application to high-throughput biology data allows re-weighting feature groups from different assays.
Conclusions:
- The novel penalized regression method effectively utilizes external covariates for adaptive penalization.
- This approach extends the applicability of penalized regression, enhances interpretability, and can boost prediction accuracy.
- The method is particularly beneficial in fields like high-throughput biology where diverse data sources are available.
Related Concept Videos
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Regression Analysis
7.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.7K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Variation
7.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.7K
Parametric Survival Analysis: Weibull and Exponential Methods
965
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
965
