Related Experiment Video
Updated: Jan 11, 2026

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
From Logistic Regression to Foundation Models: Factors Associated With Improved Forecasts
Abdulazeez Alabi1, Olajide Akinpeloye2,3, Osayimwense Izinyon4
1Mathematics and Statistics, Georgia State University, Atlanta, USA.
Abstract:
Chronic‑disease risk models using electronic health record (EHR) data inform screening and resource allocation. Calibration (expected calibration error, slope, and intercept), transportability under temporal or site shifts, and decision utility (net benefit) govern the clinical value. Narrative synthesis of comparative studies from January 2019 to October 8, 2025, appraised classical regression and gradient‑boosted decision tree (GBDT) models against deep neural networks (DNNs) and foundation backbones. Evidence indicated that modern tree-based methods often achieved lower Brier scores and external calibration errors than logistic regression, but logistic regression retained a calibration slope close to 1 under temporal drift in several datasets. DNNs frequently underestimated risk for high‑risk deciles, whereas models derived from foundation backbones improved calibration and decision utility only after local recalibration and were most efficient when labels were scarce. Across tasks, decision curves showed that net benefit increased only when recalibration maintained expected calibration error (ECE) ≤0.03. Operationally, acceptance criteria should couple the calibration slope of 0.90-1.10 with pre‑specified threshold performance and monitoring schedules.
Related Concept Videos
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

