Related Experiment Videos
Bridging Machine Learning and Clinical Endpoints: A METABRIC-Informed Simulation Study of Missing Data Imputation for
1School of Analytics & Computational Sciences, Harrisburg University of Science and Technology, Harrisburg, PA 17101, USA.
Diagnostics (Basel, Switzerland)
|June 26, 2026
Summary
Handling missing data in oncology studies significantly improves Best Overall Response (BOR) classification. Clinically informed strategies for missing data, especially progression-driven dropout, are crucial for accurate RECIST criteria assessment.
Area of Science:
- Oncology
- Biostatistics
- Machine Learning
Background:
- Missing data, particularly progression-driven dropout, biases longitudinal oncology studies and RECIST criteria response classification.
- Current machine learning imputation methods lack evaluation within clinically interpretable frameworks focused on patient-level endpoints like Best Overall Response (BOR).
Purpose of the Study:
- To propose and evaluate a clinically grounded framework for assessing imputation methods in oncology studies.
- To compare the performance of imputation strategies for Best Overall Response (BOR) classification using RECIST 1.1 criteria.
Main Methods:
- Simulated longitudinal tumor trajectories for 270 patients using Gompertz and Stein-Fojo growth models.
- Introduced realistic missingness, including progression-driven dropout, across nine follow-up visits.
- Evaluated three machine learning imputation models (LSTM, MissForest, MI) using direct and non-responder strategies, assessing performance via BOR classification metrics (accuracy, Cohen's kappa).
Main Results:
- Imputation substantially improved Best Overall Response (BOR) classification accuracy and Cohen's kappa across both simulation models.
- Non-responder imputation strategies consistently outperformed direct imputation, improving accuracy by up to 10 percentage points and kappa by up to 17 percentage points.
- Findings demonstrated robustness across tumor growth models and missingness scenarios, highlighting the impact of imputation.
Conclusions:
- Effective handling of missing data in oncology requires careful consideration of both imputation methods and clinically meaningful endpoints.
- Clinically informed strategies for missing data, particularly progression-related missingness, significantly enhance RECIST-based Best Overall Response (BOR) classification.
- Endpoint selection and estimand strategy for missing data handling may be more influential than the specific imputation model chosen.
Related Concept Videos
Kaplan-Meier Approach
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least squares (OLS)...