Related Experiment Video
Updated: Jan 4, 2026

O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
Weak signals in high-dimension regression: detection, estimation and prediction.
Yanming Li1, Hyokyoung G Hong2, S Ejaz Ahmed3
1Department of Biostatistics, University of Michigan, Ann Arbor, MI 48109 USA.
This study introduces a new method to improve statistical predictions by including weak signals often ignored by traditional techniques. This approach enhances estimation and prediction accuracy, particularly when weak signals are prevalent.
Area of Science:
- Statistics
- Econometrics
- Machine Learning
Background:
- Traditional regularization methods like Lasso, group Lasso, and SCAD prioritize strong signals, potentially leading to biased predictions when weak signals are numerous.
- Ignoring weak signals can compromise the accuracy of statistical models, especially in complex datasets.
Purpose of the Study:
- To develop a novel statistical approach that incorporates weak signals into variable selection, estimation, and prediction.
- To improve the performance of predictive models by accounting for both strong and weak effects.
Main Methods:
- A two-stage procedure involving covariance-insured screening for weak signal detection.
- Post-selection estimation using a shrinkage estimator to jointly estimate selected strong and weak signals.
- The proposed method is termed the covariance-insured screening based post-selection shrinkage estimator.
Main Results:
- The proposed method demonstrates improved estimation and prediction performance in simulation studies.
- Asymptotic properties of the new estimator have been theoretically established.
- The method was successfully applied to predict annual gross domestic product (GDP) rates.
Conclusions:
- Incorporating weak signals significantly enhances statistical estimation and prediction accuracy.
- The covariance-insured screening based post-selection shrinkage estimator offers a robust alternative to traditional methods.
- This approach has practical applications in economic forecasting and other fields requiring complex data analysis.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...

