Related Experiment Video
Updated: Feb 1, 2026

08:34
Robot-assisted Partial Splenectomy
Published on: January 2, 2026
552
Drawing inferences for high-dimensional linear models: A selection-assisted partial regression and smoothing
Zhe Fei1, Ji Zhu2, Moulinath Banerjee2
1Department of Biostatistics, University of Michigan, Ann Arbor, Michigan.
Biometrics
|December 15, 2018
Summary
This study introduces a new method for high-dimensional linear models, called Selection-assisted Partial Regression and Smoothing (SPARES). SPARES enables simultaneous estimation and inference, overcoming limitations of traditional theories and yielding meaningful genomic data analysis.
Area of Science:
- Statistics
- High-dimensional data analysis
- Genomic data analysis
Background:
- Inferences for high-dimensional models are complex due to inapplicable regular asymptotic theories.
- Existing methods may not adequately address simultaneous estimation and inference challenges.
Purpose of the Study:
- To propose a novel framework for simultaneous estimation and inference in high-dimensional linear models.
- To introduce the Selection-assisted Partial Regression and Smoothing (SPARES) procedure.
Main Methods:
- SPARES smooths partial regression estimates using variable selection to simplify to low-dimensional least squares.
- The method employs data splitting, variable selection, and partial regression.
- Asymptotic unbiasedness and normality are proven, with variance derived via a nonparametric delta method.
Main Results:
- The SPARES estimator is shown to be asymptotically unbiased and normal.
- Performance is evaluated through simulations and comparison with de-biased LASSO.
- The method successfully analyzed two genomic datasets, yielding biologically relevant findings.
Conclusions:
- SPARES provides a viable framework for simultaneous estimation and inference in high-dimensional linear models.
- The method offers a competitive alternative to existing techniques like de-biased LASSO.
- SPARES demonstrates practical utility in analyzing complex genomic data.
Keywords:
Selection-assisted Partial Regression and Smoothing (SPARES)confidence intervalshigh-dimensional inferencehypothesis testingmultisample-splittingMore Related Videos
Related Concept Videos
Regression Toward the Mean
7.0K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.0K
Multiple Regression
4.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.0K
Correlation and Regression
3.4K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.4K
Regression Analysis
8.4K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.4K
Microsoft Excel: Regression Analysis
1.6K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
1.6K
Drawing Free-body Diagrams: Rules
15.8K
The first step in describing and analyzing most phenomena in physics involves the careful drawing of a free-body diagram. Free-body diagrams are useful in analyzing forces acting on an object or system, and are employed extensively in the study and application of Newton's laws of motion. The steps to draw a free-body diagram are listed below:
15.8K

