Related Experiment Videos
Causal estimation and machine learning methods for survival outcomes among AYA cancer patients: A scoping review
Nana Owusu Mensah Essel1, Rui Liu2, Liz Dennett3
1School of Public Health, College of Health Sciences, University of Alberta, Edmonton, Alberta, Canada; Department of Emergency Medicine, Faculty of Medicine & Dentistry, College of Health Sciences, University of Alberta, Edmonton, Alberta, Canada.
Abstract:
Survivors of adolescent and young adult (AYA) cancer experience an elevated risk of premature mortality. While causal inference and machine learning (ML) methods are increasingly applied to observational survival data, the methodological landscape of this work remains unmapped. This scoping review mapped the extent, range, and nature of evidence using these methods to estimate or predict survival outcomes in this population. Following Joanna Briggs Institute (JBI) scoping review methodology, eight databases were searched from inception to November 2025. From 11,584 identified records, 5629 were screened after duplicate removal, of which 104 studies met the inclusion criteria (99 peer-reviewed journal articles and 5 conference abstracts). Of these, 68 (65.4%) were published from 2020 onward, 104 (100.0%) were retrospective cohorts, and most used US registry data (75 [72.1%]; SEER alone, 59 [56.7%]). Eighty-seven studies applied causal inference methods, among which propensity score matching was near-universal (75; 86.2%), followed by inverse probability of treatment weighting (9; 10.3%). None used targeted maximum likelihood estimation, marginal structural models for time-varying confounding, causal forests, or target trial emulation. Methods commonly associated with causal inference were widely used, but formally specified causal analyses were uncommon. Only three studies (3.4%) were designed causal analyses (regression discontinuity, causal mediation, and inverse probability weighted analysis [1 study each]), one (1.1%) provided a complete identification statement, and one (1.1%) reported a sensitivity analysis for unmeasured confounding. Competing risks were relevant where a quantity of specific event(s) was estimated and death from other causes could prevent the observation of specific events. Among the 48 applicable causal studies, 9 (18.8%) used a competing risks or relative survival method. ML prediction methods were used in nineteen studies, most often random survival forests (12; 63.2%); only four (21.1%) were externally validated. Among 99 peer-reviewed studies, the most common methodological issues were competing risks mishandling (22/99; 22.2% overall) and model overfitting (15/99; 15.2%). Research applying causal inference and ML methods to survival outcomes in AYA cancer patients is growing rapidly. However, the evidence draws on a narrow range of methods and data, dominated by propensity score matching using US registry data, with minimal documentation of identifying assumptions, rare probing of unmeasured confounding, and competing risks rarely addressed where the outcome made them relevant. ML models showed adequate discrimination, but a large proportion were not externally validated or calibrated, raising concerns about external validity and generalizability. Closing the gap requires adopting target trial-emulated causal designs with estimand-appropriate competing risk handling, together with externally validated, calibrated mortality prediction tools.
Related Concept Videos
Cancer Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...