Related Experiment Video
Updated: Dec 27, 2025

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.9K
High-dimensional regression in practice: an empirical study of finite-sample prediction, variable selection and
Fan Wang1, Sach Mukherjee2, Sylvia Richardson1
11MRC Biostatistics Unit, University of Cambridge, Cambridge, UK.
Summary
No single penalized regression method excels in all high-dimensional settings. This large-scale comparison reveals performance variations across goals like prediction and variable selection, offering practical guidance for method selection.
Area of Science:
- Statistics
- Machine Learning
- Computational Biology
Background:
- Penalized likelihood methods are crucial for high-dimensional regression analysis.
- Existing theory is well-developed, but practical performance in finite samples is unclear.
- Empirical studies are needed to guide users in selecting appropriate methods.
Purpose of the Study:
- To conduct a large-scale empirical comparison of various penalized regression methods.
- To evaluate method performance for prediction, variable selection, and variable ranking.
- To provide practical insights into method suitability based on data characteristics.
Main Methods:
- Comparison of Lasso, Adaptive Lasso, Elastic Net, Ridge Regression, SCAD, Dantzig Selector, and Stability Selection.
- Utilized over 2300 synthetic and semi-synthetic data-generating scenarios.
- Systematically varied factors including sample size, dimensionality, sparsity, signal strength, and multicollinearity.
Main Results:
- Significant performance variations observed among different penalized regression methods.
- No single method demonstrated superior performance across all scenarios and goals.
- Results indicate that method choice depends on the specific objective and data properties.
Conclusions:
- A 'no panacea' approach is supported; method selection requires careful consideration of the task and data.
- The study provides empirical evidence to complement theoretical understanding.
- Offers practical recommendations for choosing penalized regression techniques in high-dimensional settings.
Related Concept Videos
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Friedman Two-way Analysis of Variance by Ranks
439
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
439
Survival Tree
335
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
335
Ranks
415
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
415
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Variation
7.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
7.7K

