Related Experiment Videos
Understanding palliative care trajectories through clustering, calibrated prediction models, and explainable AI
Battushig Migiddorj1, Marijka Batterham2, David Currow3
1School of Computing and Information Technology, University of Wollongong, New South Wales, Australia.
Background:
Despite growing demand for palliative care, access remains limited and often occurs late in life. Existing machine learning approaches rarely integrate clinically validated measures of symptom burden and functional status with healthcare utilization, limiting the interpretability of care trajectories.
Objectives:
To (1) identify distinct patient subgroups based on symptom burden and functional status, and (2) predict palliative care episode duration using the identified subgroups.
Method:
A retrospective cohort study was conducted using national data from the Australian Palliative Care Outcomes Collaboration, including adults (≥18 years) who died between 2014 and 2023 (261,290 patients). Unsupervised clustering (Partitioning Around Medoids, K-means, Hierarchical, Gaussian Mixture Models) was used to derive symptom-function clusters with bootstrap stability and clinical validity assessed. Supervised prediction models (Elastic Net, Random Forest, XGBoost) predicted care duration using out-of-fold validation and a held-out test set. Post-hoc calibration addressed right-skewed outcomes. Model performance and prediction error were examined across diagnostic groups, clusters, and clinically defined episode duration categories. Explainability analyses were conducted to interpret model behavior.
Results:
Clustering consistently identified a severity gradient, with a critical transition in mid-range functional states where symptom burden and care needs escalated rapidly. Predictive performance was modest (R² ≈ 0.2 after calibration), reflecting the complexity of palliative care trajectories. Calibration improved agreement between predicted and observed outcomes, particularly for longer episodes. In contrast, clustering did not improve predictive accuracy but enhanced interpretability by summarizing patient heterogeneity. Prediction accuracy varied systematically across clusters, diagnostic groups, and episode duration, with higher accuracy in shorter, late-stage episodes and greater uncertainty in longer trajectories. Functional status and episode type were the dominant drivers of care duration.
Conclusions:
Palliative care needs follow a continuous severity gradient, with prediction accuracy varying across disease stages, most closely aligned to functional status. Prediction is most accurate near the end of life and least accurate earlier, highlighting a clinical paradox where intervention potential is greatest when predictions are most uncertain. Methodologically, these findings show that calibration improves prediction reliability, while clustering enhances interpretability without increasing accuracy. The value of machine learning lies not in maximizing accuracy alone, but in combining calibrated predictions with interpretable stratification to support stage-specific care planning.