Related Experiment Video
Updated: Sep 10, 2025

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Mastering rare event analysis: subsample-size determination in Cox and logistic regressions
Tal Agassi1, Nir Keret1, Malka Gorfine1
1Department of Statistics and Operations Research, Tel Aviv University, Tel Aviv 69978, Israel.
This study introduces tools for selecting optimal subsample sizes in massive data analysis, improving efficiency for Cox regression and logistic regression models, especially with imbalanced data.
Area of Science:
- Computational Statistics and Big Data Analytics.
- Methodological frameworks for subsample-size determination in high-dimensional modeling.
- Intersection of survival analysis and rare event analysis in massive datasets.
Background:
Modern data science frequently encounters massive datasets that impose significant burdens on computational memory and processing time. Prior research has shown that optimal subsampling methods can reduce these burdens while minimizing the loss of statistical efficiency. These existing frameworks allow researchers to analyze smaller portions of a dataset to approximate the parameters of the full population. However, current methodologies often fail to provide a rigorous mathematical basis for deciding exactly how many observations are required for a specific analysis. This lack of guidance is particularly problematic when dealing with rare events where data points of interest are scarce. Statistical practitioners currently rely on heuristic approaches or arbitrary thresholds that may not preserve the necessary power for complex regressions. This absence of evidence motivated the development of formal tools to guide the selection of appropriate sample sizes within subsampling routines.
Purpose Of The Study:
This research develops specific tools for determining the optimal subsample size across three distinct statistical modeling environments. The investigators focus on the Proportional Hazards (Cox) regression model for survival data characterized by infrequent occurrences. The study also addresses Logistic Regression (LR) within the context of both balanced and imbalanced datasets to ensure broad applicability. A secondary objective involves creating a novel optimal subsampling procedure specifically designed for imbalanced logistic data. The authors aim to bridge the existing gap between theoretical subsampling efficiency and practical implementation in large-scale studies. By providing these selection tools, the work seeks to optimize the trade-off between computational speed and statistical precision. The project intends to validate these new methodologies through both synthetic simulations and real-world massive data applications.
Main Methods:
The researchers implemented an extensive simulation study to evaluate the performance of the proposed size-selection tools. They applied the Proportional Hazards (Cox) regression framework to a massive dataset from the UK Biobank containing colorectal cancer records. This specific survival analysis involved processing approximately 350 million rows of data to test the limits of the subsampling algorithm. For the Logistic Regression (LR) component, the team utilized linked birth and infant death records comprising roughly 28 million observations. The methodology incorporated a new optimal subsampling procedure tailored for imbalanced datasets where the event of interest is rare. Statistical comparisons focused on the minimized efficiency loss achieved by the selected subsample sizes relative to full-dataset benchmarks. The analytical framework prioritized the maintenance of predictive accuracy while drastically reducing the required computational memory.
Main Results:
The simulation study confirmed that the new tools effectively identify subsample sizes that maintain high statistical efficiency. Analysis of the UK Biobank colorectal cancer data demonstrated that the Cox regression model could be accurately estimated using the determined subsample. Results from the 28 million linked birth and infant death observations showed that the imbalanced logistic regression procedure performed reliably. The proposed tools successfully balanced the computational time requirements with the need for precise parameter estimation in rare event scenarios. The optimal subsampling procedure for imbalanced data yielded lower variance in coefficient estimates compared to standard random sampling. These findings suggest that the size-selection tools provide a robust solution for managing massive datasets without sacrificing analytical rigor. The empirical evidence supports the use of these tools across diverse data structures, including those with extreme class imbalances.
Conclusions:
The development of these size-selection tools provides a necessary foundation for efficient large-scale data analysis in modern research. These findings imply that researchers can now make informed decisions regarding the volume of data needed for Cox and logistic regressions. The application to UK Biobank colorectal cancer data highlights the potential for these methods to accelerate epidemiological studies. Future investigations might apply these subsampling principles to other complex models beyond survival and logistic frameworks. The study establishes that imbalanced datasets require specialized subsampling procedures to ensure the accurate detection of rare events. Implementation of these tools could significantly reduce the environmental and financial costs associated with high-performance computing. The authors conclude that this framework represents a significant advancement in the practical utility of optimal subsampling for massive datasets.
Frequently Asked Questions
The tools identify the minimum number of observations required to approximate full-dataset parameters. In the UK Biobank colorectal cancer analysis, this mechanism ensured that the Proportional Hazards (Cox) model maintained high statistical efficiency while processing a fraction of the original 350 million rows.
The researchers utilized a linked birth and infant death dataset containing approximately 28 million observations. This massive scale allowed the team to demonstrate that their optimal subsampling procedure for imbalanced data could accurately estimate parameters even when the target event is extremely rare.
This dataset was chosen because its massive size of 350 million rows presents significant computational challenges for standard survival analysis. Using this model system allowed the authors to prove that their size-selection tools can handle real-world big data with rare survival events.
The study focuses specifically on the Proportional Hazards (Cox) regression model for survival data and Logistic Regression (LR) for balanced or imbalanced datasets. The authors have not yet extended these specific tools to other non-linear models or different types of longitudinal data analysis.
The study's authors propose that these tools bridge the gap between theoretical subsampling and practical implementation. They conclude that using these procedures will allow researchers to manage massive datasets more judiciously, reducing computational time and memory demands without losing significant statistical precision.
Related Concept Videos
Censoring Survival Data
The Mantel-Cox Log-Rank Test
Assumptions of Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Survival Tree
Building a Survival Tree
Constructing a...

