Related Experiment Videos
Bridging Machine Learning and Clinical Endpoints: A METABRIC-Informed Simulation Study of Missing Data Imputation for
1School of Analytics & Computational Sciences, Harrisburg University of Science and Technology, Harrisburg, PA 17101, USA.
None:
Background: Missing data, particularly progression-driven dropout, introduces substantial bias in longitudinal oncology studies, directly impacting response classification based on RECIST criteria. While machine learning-based imputation methods are increasingly used, their performance is rarely evaluated in a clinically interpretable framework centered on patient-level endpoints such as Best Overall Response (BOR). Methods: We propose a clinically grounded evaluation framework based on RECIST 1.1 focused on patient-level Best Overall Response classification. Longitudinal tumor trajectories were simulated for 270 patients (1:1, HER2+ and HER2-) across nine follow-up visits using both Gompertz and Stein-Fojo growth models, resulting in 2700 patient-visit observations. Realistic missingness was introduced through a combination of random mechanisms and progression-driven dropout. Three machine learning imputation models, long short-term memory (LSTM), MissForest, and Multiple Imputation (MI) were evaluated under both direct (MAR-based) and non-responder imputation strategies. Performance was assessed using BOR classification metrics, including accuracy and Cohen's kappa. Result: Across both simulation frameworks, imputation substantially improved BOR classification performance. Under the Gompertz model, accuracy increased from 0.84-0.89 with direct imputation to 0.94-0.99 with non-responder imputation, with corresponding kappa improvements from 0.73-0.82 to 0.90-0.99. Similar trends were observed under the Stein-Fojo model (accuracy: 0.82-0.84 vs. 0.91-0.96; kappa: 0.69-0.72 vs. 0.86-0.94). Across all evaluated methods, NRI improved classification performance by approximately 10 percentage points in accuracy and up to 17 percentage points in kappa. The improvement was observed consistently across both tumor growth models and different missingness scenarios, demonstrating the robustness of the findings. Conclusions: This study demonstrates that successful handling of missing data depends not only on the imputation method itself, but also on the choice of a clinically meaningful endpoint and appropriate estimand strategies aligned with the underlying missing data assumptions. In the METABRIC-derived simulations, clinically informed handling of progression-related missingness substantially improved RECIST-based BOR classification across all evaluated methods, suggesting that appropriate endpoint selection and the corresponding estimand strategy for missing data handling may have a greater influence on classification performance than the choice among the imputation models applied.
Related Concept Videos
Kaplan-Meier Approach
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Mechanistic Models: Compartment Models in Individual and Population Analysis