Related Experiment Video
Updated: Jan 17, 2026

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.2K
Internal validation strategy for high dimensional prognosis model: A simulation study and application to
Antoine Dubray-Vautrin1,2, Victor Gravrand3, Grégoire Marret4
1Institut Curie, INSERM, Saint Cloud U1331, France.
Computational and Structural Biotechnology Journal
|September 24, 2025
Summary
For high-dimensional oncology models, k-fold cross-validation and nested cross-validation are recommended for internal validation. These methods offer more stable and reliable performance than train-test or bootstrap approaches, especially with adequate sample sizes.
Area of Science:
- Oncology
- Bioinformatics
- Statistical Modeling
Background:
- High-dimensional data, including genomics and transcriptomics, are increasingly vital for predictive oncology models forecasting time-to-event outcomes.
- Internal validation is essential to reduce optimism bias before external validation of these predictive models.
- Existing validation strategies like train-test, bootstrap, and cross-validation lack established benchmarks in high-dimensional contexts.
Purpose of the Study:
- To compare the performance of common internal validation strategies for high-dimensional predictive models.
- To provide evidence-based recommendations for internal validation methods in transcriptomic analysis for oncology.
- To evaluate the stability and reliability of train-test, bootstrap, k-fold cross-validation, and nested cross-validation.
Main Methods:
- A simulation study utilized head and neck cancer cohort data (SCANDARE) with clinical and transcriptomic variables, simulating disease-free survival.
- Datasets were simulated across various sample sizes (50 to 1000) with 100 replicates each.
- Cox penalized regression was employed for model selection, followed by train-test, bootstrap, 5-fold cross-validation, and nested cross-validation (5x5) to assess performance using time-dependent AUC, C-Index, and integrated Brier Score.
Main Results:
- Train-test validation exhibited unstable performance across simulations.
- Bootstrap methods showed over-optimistic (conventional) or overly pessimistic (0.632+) results, particularly with smaller sample sizes (n=50-100).
- K-fold and nested cross-validation demonstrated improved and more stable performance with increasing sample sizes, though nested cross-validation's stability varied with regularization methods.
Conclusions:
- K-fold cross-validation and nested cross-validation are recommended for internal validation of Cox penalized models in high-dimensional, time-to-event oncology settings.
- These cross-validation techniques offer superior stability and reliability compared to train-test or bootstrap methods, especially when sufficient sample sizes are available.
- The choice between k-fold and nested cross-validation may depend on the specific regularization techniques used in model development.

