Related Experiment Video
Updated: May 13, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Sample size considerations of prediction-validation methods in high-dimensional data for survival outcomes
1Department of Biostatistics and Bioinformatics, Duke University School of Medicine, Durham, North Carolina, USA.
Genetic Epidemiology
|March 9, 2013
Summary
10-fold cross-validation with permutation is the best strategy for validating genome data prediction models with survival outcomes. This study also introduces a sample size calculation method for designing such validation studies.
Area of Science:
- Genomics
- Biostatistics
- Clinical Prediction Modeling
Background:
- High-dimensional genomic data is frequently used to build prediction models for clinical outcomes.
- Model validation is crucial but the performance of various validation methods, particularly for survival outcomes, lacks comprehensive investigation.
- Existing validation methods need rigorous evaluation to enhance statistical power and control Type I error.
Purpose of the Study:
- To comprehensively investigate the statistical properties and performance of different validation methods for prediction models using high-dimensional genomic data and survival outcomes.
- To identify the most effective validation strategy for controlling Type I error and maximizing statistical power.
- To develop a sample size calculation strategy to aid in the design of future validation studies.
Main Methods:
- Extensive simulations were conducted to examine the statistical properties of various validation strategies.
- The performance of validation methods was evaluated using both simulated data and a real-world clinical dataset.
- 10-fold cross-validation combined with permutation testing was specifically analyzed.
Main Results:
- 10-fold cross-validation with permutation demonstrated superior statistical power while maintaining Type I error rates close to the nominal level.
- The study identified this method as the most effective for validating prediction models with survival outcomes.
- A novel sample size calculation method was developed based on the findings.
Conclusions:
- 10-fold cross-validation with permutation is recommended as the optimal validation strategy for high-dimensional genomic data and survival outcomes.
- The developed sample size calculation method can facilitate the design of robust validation studies in biomedical research.
- This research provides a framework for more powerful and reliable validation of predictive models in genomics.
Related Concept Videos
Survival Tree
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a survival tree begins...
Building a Survival Tree
Constructing a survival tree begins...
Assumptions of Survival Analysis
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
Kaplan-Meier Approach
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
Comparing the Survival Analysis of Two or More Groups
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and Cox...
Parametric Survival Analysis: Weibull and Exponential Methods
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Censoring Survival Data
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different reasons...
