Related Experiment Video
Updated: Feb 20, 2026

04:57
Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
10.9K
A new framework for prediction and variable selection for uncommon events in a large prospective cohort study
Hye-Seung Lee1, Jeffrey P Krischer1
1Health Informatics Institute, 3650 Spectrum Blvd., Suite 100, University of South Florida, Tampa, Florida 33612.
Summary
Predicting rare events in large studies is challenging. A nested case-control design offers a powerful alternative to full cohort analysis, improving statistical power for prediction model validation.
Area of Science:
- Epidemiology
- Biostatistics
- Medical Informatics
Background:
- Traditional data splitting for prediction model validation reduces statistical power when predicting uncommon events.
- Large prospective cohort studies present unique challenges for validating prediction models of rare outcomes.
Purpose of the Study:
- To propose and illustrate a nested case-control design for validating prediction models of uncommon events in large prospective cohort studies.
- To introduce a framework for iterative variable selection and missing data imputation within this design.
Main Methods:
- Utilizing a nested case-control design within a large prospective cohort study.
- Implementing iterative variable selection with random forest for missing data imputation.
- Fitting a prediction model on selected variables within the case-control cohort.
- Validating the model by comparing specificity in the case-control cohort versus the remaining cohort.
Main Results:
- The nested case-control design provides a valid alternative for prediction model validation in large cohorts, particularly for rare events.
- The proposed iterative variable selection and imputation methods enhance model robustness.
- The framework demonstrated effective application in high-dimensional variable selection.
Conclusions:
- A nested case-control design is a statistically efficient approach for validating prediction models of rare events in large prospective cohorts.
- The proposed methodology addresses challenges in variable selection and missing data, enabling robust model validation.
- This framework facilitates accurate prediction model development and validation in epidemiological research.
Related Concept Videos
Prediction Intervals
3.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.5K
Survival Tree
440
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
440
Study Design in Statistics
10.1K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
10.1K
Introduction To Survival Analysis
862
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
862
Assumptions of Survival Analysis
448
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
448
Statistical Methods for Analyzing Epidemiological Data
1.0K
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
1.0K

