Related Experiment Videos
Infant mortality across EU health systems: heterogeneity encoding, fixed effects regression, and machine learning in
1Department of Technology, Faculty of Health Sciences and Technology, Kristiania University of Applied Sciences, Oslo, Norway.
Introduction:
The infant mortality rate is widely used in comparative health research, but less is known about how panel regression and machine-learning models perform in small macro-level health datasets when they are evaluated using comparable information and the same out-of-period test period.
Methods:
Using a balanced panel of 25 European Union countries observed annually from 2000 to 2022 (575 country-year observations, no missing data) assembled from OECD and World Bank sources, we separated an explanatory analysis from a predictive analysis. The explanatory layer used a two-way fixed effects regression with country-clustered standard errors. The predictive layer trained pooled linear regression, tuned Random Forest, Gradient Boosting, and a Multilayer Perceptron, with and without explicit country information, on 2000-2017 and evaluated them on a held-out 2018-2022 window, alongside two naive benchmarks (persistence and historical country mean) and two prediction-compatible fixed effects specifications. Hyperparameters were tuned by expanding-window temporal cross-validation within the training period, leaving the test window untouched.
Results:
Under common out-of-period evaluation, a simple persistence benchmark achieved the strongest predictive performance (R2 = 0.74), closely followed by tuned Gradient Boosting and Random Forest with country information (R2 = 0.71 and 0.69). Explicit country encoding improved the tree-based models, but adding country information was not enough to improve the linear specifications substantially. The prediction-compatible fixed effects and pooled linear models both produced negative out-of-period R2 values. Model rankings varied across evaluation windows: in an alternative pre-pandemic split, tuned nonlinear models outperformed persistence. In the explanatory fixed effects model, GDP per capita and the total fertility rate were significant under conventional country-clustered inference, although neither remained significant under a small-cluster wild-bootstrap sensitivity analysis; the substantive covariate block had a partial R2 of 0.366 net of country and year effects.
Conclusion:
In this EU infant-mortality panel, much of the predictive signal came from persistent country-specific outcome history. Country information improved the tree-based models, while model flexibility also remained important. The strong within-panel fit of the fixed-effects model did not translate into strong out-of-period prediction.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Bias in Epidemiological Studies
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...
Regression Toward the Mean
Comparing the Survival Analysis of Two or More Groups
Longitudinal Studies