Related Experiment Video
Updated: Dec 11, 2025

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Missing data imputation and synthetic data simulation through modeling graphical probabilistic dependencies between
Mireia Vilardell1, Maria Buxó2, Ramon Clèries3
1Sección de Estadística del Departamento de Genética, Microbiología y Estadística de la Facultad de Biología, Universidad de Barcelona, 08028, Spain.
Graphical modeling (GM) improves breast cancer survival analysis by addressing missing data and small sample sizes. GM models outperform standard methods for imputation and synthetic data generation, enhancing predictive accuracy.
Area of Science:
- Biostatistics
- Cancer Research
- Machine Learning
Background:
- Population-based breast cancer (BC) survival studies face challenges with missing predictive variables (e.g., Stage) and small sample sizes due to class imbalance.
- These issues necessitate advanced data modeling and simulation techniques for reliable analysis.
Purpose of the Study:
- To introduce ModGraProDep, a graphical modeling (GM) procedure designed to address missing data and small sample sizes in BC survival studies.
- To compare the performance of GM-derived models against conventional classification, machine learning, and oversampling algorithms.
Main Methods:
- Developed and applied the ModGraProDep procedure using graphical modeling on two validated BC datasets.
- Assessed performance in scenarios with data missing completely at random (MCAR) and missing not at random (MNAR).
- Compared GM models (GM.SAT, GM.K1, GM.TEST) against standard algorithms for missing data imputation and synthetic data simulation.
Main Results:
- GM models, particularly GM.K1 and GM.TEST, demonstrated superior prediction performance compared to other algorithms in both MCAR and MNAR scenarios.
- While GM.SAT showed strong performance, its direct predictions might yield unreliable conclusions; however, synthetic data from it was the poorest strategy.
- The remaining GM models offered a better alternative to oversampling for synthetic data generation.
Conclusions:
- The presented GM procedure is recommended for one-variable imputation/prediction of missing data in BC survival studies.
- GM-derived synthetic datasets can effectively augment small or imbalanced datasets, proving valuable for clinical applications like predictive risk analysis.
Related Concept Videos
Assumptions of Survival Analysis
Cancer Survival Analysis
Survival Tree
Building a Survival Tree
Constructing a...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Kaplan-Meier Approach
Censoring Survival Data

