在低维数据中,在完全和选择的线性回归模型中进行估计后收缩,并进行了重新检查
Edwin Kipruto1, Willi Sauerbrei1
1Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center - University of Freiburg, Freiburg, Germany.
Biometrical journal. Biometrische Zeitschrift
|September 27, 2024
概括
估计后收缩方法可以提高回归模型预测的准确性. 非负参数智能收缩 (NPWS) 在完整模型中表现出色,而处罚方法在高相关性和低信号噪声比等具有挑战性的条件中优越.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
背景情况:
- 过度装配会降低对新数据的回归模型性能.
- 变量选择可以在回归估计中引入偏差.
- 收缩方法可以减轻过和偏差.
研究的目的:
- 评估估计后收缩,以提高预测性能.
- 将收缩方法与普通最小平方 (OLS),,最佳子集选择 (BSS) 和拉索相比较.
- 引入和评估一种新的非负参数智能收缩 (NPWS) 方法.
主要方法:
- 模拟研究比较预测错误和变量选择.
- 使用OLS,和收缩方法对完整模型的评估.
- 评估使用BSS,拉索和估计后收缩的选定模型.
主要成果:
- 在完整模型中,NPWS的表现优于全球收缩;PWS的表现低于OLS.
- 在低相关性/高SNR方面,NPWS优于峰;在小样本/高相关性/低SNR方面,峰表现最好.
- 在选定的模型中,估计后收缩方法的表现类似,全球收缩略低一些.
- 拉索在特定条件下 (小样本,低SNR,高相关性) 优于BSS和估计后收缩.
结论:
- 当有足够的数据可用时,NPWS可以提高预测准确度,而不是全球收缩.
- 处罚方法在高相关性,小样本大小和低SNR场景中通常优于估计后收缩.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Regression Analysis
5.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.6K
Truncation in Survival Analysis
173
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
173
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
411
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
411


