Machine Learning for Predicting 12-Month Survival after Liver Transplantation: Temporal Validation, Calibration, and
Jose Ricardo de Oliveira1, Adriano Galindo Leal1, Rodrigo Luiz Macacari2
1Artificial Intelligence and Analytics Department, Institute for Technological Research, São Paulo 05508-901, SP, Brazil.
Abstract:
Accurate prediction of posttransplant survival is critical for optimizing liver allocation under organ scarcity. Although the Model for End-Stage Liver Disease (MELD) is widely used for prioritization, it was not designed to estimate posttransplant outcomes, limiting its utility for individualized risk stratification. We developed and temporally validated machine-learning models to predict 12-month survival after liver transplantation using routinely collected donor-recipient variables from a real-world cohort. Sequential temporal validation was used to approximate prospective deployment. Performance was evaluated in terms of discrimination (receiver operating characteristic-area under the curve [ROC-AUC] and precision-recall area under the curve), calibration (Brier score), and decision-analytic evaluation using decision curve analysis. Class imbalance was addressed through cost-sensitive learning, and predicted probabilities were recalibrated with Platt scaling. Across temporal validation windows, discrimination remained modest, with ROC-AUC values below 0.70 and reaching 0.599 in the final window, reflecting the complexity of posttransplant outcomes in heterogeneous clinical populations. Calibration remained relatively stable (Brier score ≈ 0.18). Within the evaluated threshold range, the machine-learning models yielded a positive net benefit and generally exceeded the transformed MELD baseline in decision-analytic comparisons. However, these findings should be interpreted cautiously given the modest discrimination and the exploratory nature of the MELD transformation used for cross-model comparison. Logistic regression provided the most consistent balance of discrimination, calibration, and net benefit. Overall, calibrated models with modest discrimination may still offer complementary risk information. However, further external validation, threshold-specific evaluation, and prospective testing are required before such models can be considered for operational decision support.
