从后勤回归到基础模型:与改善预测相关的因素
Abdulazeez Alabi1, Olajide Akinpeloye2,3, Osayimwense Izinyon4
1Mathematics and Statistics, Georgia State University, Atlanta, USA.
Cureus
|November 17, 2025
概括
针对慢性疾病风险的电子健康记录 (EHR) 模型需要仔细校准. 现代方法有希望,但物流回归保持稳定性,再校准是临床实用性的关键.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的机器学习
- 临床流行病学临床流行病学
背景情况:
- 电子健康记录 (EHR) 对慢性疾病风险建模至关重要,影响查和资源分配.
- 这些模型的临床实用性取决于校准,可运输性和决策实用性 (净效益).
研究的目的:
- 为了比较评估经典回归,渐变增强决策树 (GBDT),深度神经网络 (DNN) 和慢性疾病风险预测的基础骨干.
- 评估模型在校准,可运输性和决策实用性方面的性能.
主要方法:
- 2019年1月至2025年10月期间发表的比较研究的叙事综合.
- 对物流回归,GBDT,DNN和基础骨干模型的评估,使用比如Brier分数,校准误差 (斜率,截取) 和决策曲线 (净效益) 等指标.
主要成果:
- 现代基于树的方法 (GBDTs) 在布里尔分数和外部校准错误中经常超过后勤回归.
- 后勤回归证明了在时间漂移下优越的校准斜率稳定性.
- 深度神经网络 (DNN) 倾向于低估高风险群体中的风险;基础骨干需要重新校准以提高性能.
- 对模型来说,重新校准对于实现净收益至关重要,特别是当预期校准误差 (ECE) 保持在0.03.以下时.
结论:
- 虽然先进的模型具有潜力,但物流回归显示校准斜率的稳定性.
- 有效的重新校准策略对于提高风险预测模型的临床效用和决策至关重要.
- 运行验收标准应整合校准斜率 (0.90-1.10),性能值和监测时间表.
相关概念视频
Residuals and Least-Squares Property
9.0K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.0K
Regression Analysis
7.9K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.9K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Parametric Survival Analysis: Weibull and Exponential Methods
1.0K
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
1.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
271
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
271


