Multivariate survival analysis in big data: A divide-and-combine approach
Wei Wang1, Shou-En Lu1, Jerry Q Cheng2
1Department of Biostatistics and Epidemiology, Rutgers University, Piscataway, New Jersey, USA.
Biometrics
|April 13, 2021
Summary
Analyzing large-scale multivariate failure time data is now feasible with a novel divide-and-combine method. This approach overcomes computational challenges in big data, offering accurate risk factor identification for health outcomes.
Area of Science:
- Biostatistics
- Health Informatics
- Epidemiology
Background:
- Multivariate failure time data analysis commonly uses marginal proportional hazards or frailty models.
- Extremely large datasets present significant computational challenges for traditional analysis methods.
Purpose of the Study:
- To propose a divide-and-combine method for analyzing large-scale multivariate failure time data.
- To address computational limitations encountered with massive datasets.
Main Methods:
- A divide-and-combine strategy is proposed, involving random data subsetting and weighted estimator combination.
- Regularized estimation is applied to the combined estimator for risk factor screening.
- Theoretical properties including consistency and oracle properties are investigated.
Main Results:
- The proposed combined estimator is asymptotically equivalent to the full data estimator.
- The method effectively handles large-scale data, as demonstrated by application to the Myocardial Infarction Data Acquisition System (MIDAS) dataset.
- Simulation studies confirm the method's performance.
Conclusions:
- The divide-and-combine method provides a computationally efficient and statistically sound approach for large-scale multivariate failure time data analysis.
- This method facilitates the identification of risk factors for complex health outcomes in massive datasets.
Keywords:
big dataconfidence distributiondivide and combinemarginal modelmultivariate failure timeproportional hazards modelregularizationvariable selectionMore Related Videos
Related Concept Videos
Survival Tree
210
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
210
Comparing the Survival Analysis of Two or More Groups
384
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
384
Cancer Survival Analysis
504
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
504
Introduction To Survival Analysis
478
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
478
Assumptions of Survival Analysis
234
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
234
Parametric Survival Analysis: Weibull and Exponential Methods
793
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
793


