Related Experiment Videos
Development and Temporal External Validation of a Parsimonious, Interpretable Machine Learning Model for Predicting
Yun-Cheng Tsai1, Yi-Lin Wu2, Yung-Chen Yu2
1PecuLab LLC, Seattle, WA, United States.
Background:
Early mortality after long-term care facility (LTCF) admission is common; yet, prognostic tools are often derived from Western minimum dataset-based cohorts or require hospital electronic health record linkages that are unavailable at intake in many LTCFs. There is also limited evidence on explainable, admission-feasible machine learning-based prognostication in Asian LTCF settings.
Objective:
The aim of the study is to develop and temporally externally validate an interpretable machine learning model for predicting 6-month all-cause mortality among older adults newly admitted to LTCFs in Taiwan using routinely collected LTCF assessment data.
Methods:
We conducted a retrospective cohort study using the JUBO Long-Term Care Database, a nationwide private administrative registry covering 636 LTCFs (37.4% of the national LTCFs) in Taiwan. We included residents with first-time LTCF admission and prespecified nonoverlapping cohorts for temporal validation: development (January 1, 2020, to December 31, 2023; n=23,901) and external validation (January 1 to December 31, 2024; n=6216). The outcome measure was death within 180 days of admission. We compared a nonlinear ensemble model (hybrid of extreme gradient boosting and random forest [HybridXGBRF]) with 7 other algorithms, including tree-based and linear benchmarks. Discrimination (area under the receiver operating characteristic curve [AUROC]), classification metrics (accuracy, precision, recall, and F1), and calibration (Brier score and calibration plots) were assessed. Model interpretability was examined using Shapley Additive Explanations.
Results:
In the development cohort, 5272 of 23,901 (22.1%) residents died within 180 days. In the 2024 temporal validation cohort, 1781 of 6216 (28.7%) residents died. In internal cross-validation, HybridXGBRF had the highest AUROC among the evaluated models (0.89, 95% CI 0.88-0.89). In temporal validation, HybridXGBRF maintained strong discrimination (AUROC 0.90, 95% CI 0.89-0.91), with an accuracy of 0.85 and an F1-score of 0.68. Calibration plots indicated close agreement between predicted and observed risks across most probability ranges, with mild divergences at higher predicted risks. Shapley Additive Explanations analysis identified frequent hospitalizations within 6 months, activities of daily living impairment, and weight loss as influential predictors. The model showed stable AUROC across sex and age strata (0.89-0.90) and maintained high discrimination among residents with improving activities of daily living scores (AUROC 0.91, 95% CI 0.90-0.92).
Conclusions:
An interpretable machine learning model using routinely collected Taiwanese LTCF assessment data achieved strong discrimination, acceptable calibration, and stable temporal validation performance without requiring hospital-based electronic health record linkage. The HybridXGBRF had the highest AUROC among the evaluated models, but performance differences from extreme gradient boosting were small. The model may serve as a risk-stratification tool to help identify residents who could benefit from structured care review, advance care planning, or palliative care assessment. Prospective implementation studies could determine its impact on care processes and resident- or family-centered outcomes.