通过整合来自异质群体的外部信息来改善线性回归模型的预测:詹姆斯-斯坦估计器
Peisong Han1, Haoyue Li2, Sung Kyun Park3
1Biostatistics Innovation Group, Gilead Sciences, 333 Lakeside Drive, Foster City, CA 94404, United States.
Biometrics
|August 5, 2024
概括
本研究将外部线性回归模型总结整合到内部模型中,以提高预测准确度. 适应的詹姆斯-斯坦收缩方法提高了预测性能,即使人口异质.
科学领域:
- 生物统计学 生物统计学
- 流行病学 流行病学
- 统计建模 统计建模
背景情况:
- 内部研究使用个人数据构建线性回归模型.
- 外部研究提供了减少模型系数估计,而没有个人数据.
- 在不同的研究群体中存在异质性.
研究的目的:
- 将外部模型总结信息集成到内部模型中.
- 通过利用外部数据来提高预测准确度.
- 为了应对人口异质性带来的挑战.
主要方法:
- 适应詹姆斯-斯坦收缩方法.
- 开发用于信息整合的新型估计器.
- 模拟研究用于评估估计器性能.
主要成果:
- 拟议的估计器证明了更好的预测平均平方误差.
- 无论人口异质性如何,有效性都保持不变.
- 对骨度预测模型的成功应用.
结论:
- 詹姆斯-斯坦收缩适应有效地整合了外部模型总结.
- 这种方法在异质性存在的情况下提高了预测准确性.
- 该方法为元分析和数据集成提供了有价值的工具.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
444
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
444
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Improving Translational Accuracy
9.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.6K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K


