Related Experiment Video
Updated: Jul 10, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
A comparison of several methods for analyzing censored data.
1Exposure Assessment Solutions, Inc., Morgantown, West Virginia, USA. phewett_2006_07@oesh.com
The Annals of Occupational Hygiene
|October 18, 2007
Summary
Maximum Likelihood Estimation (MLE) and log-probit regression (LPR) methods performed best for analyzing censored occupational exposure data. Standard MLE was superior overall, especially for estimating the 95th percentile and mean, outperforming non-parametric methods.
Area of Science:
- Occupational Health
- Environmental Science
- Biostatistics
Background:
- Censored data, common in occupational exposure assessments, presents challenges for statistical analysis.
- Accurate estimation of exposure percentiles and means is crucial for risk assessment and regulatory compliance.
- Existing methods for handling data below the limit-of-detection (LOD) vary in performance.
Purpose of the Study:
- To compare the statistical performance of various methods for analyzing censored occupational exposure data.
- To evaluate methods for estimating the 95th percentile and mean of right-skewed exposure data.
- To assess method robustness under different scenarios, including varying sample sizes, LODs, and distribution types.
Main Methods:
- Evaluated maximum likelihood estimation (MLE), log-probit regression (LPR), substitution methods, non-parametric (NP) quantile methods, and Kaplan-Meier (KM) method.
- Utilized computer-generated censored datasets across diverse scenarios: varying geometric standard deviation, LOD, and sample size.
- Assessed performance using bias and root mean square error (rMSE) for estimating the 95th percentile and mean.
Main Results:
- No single method excelled in all scenarios; MLE and LPR-based methods showed consistent performance across most conditions.
- MLE methods generally yielded lower rMSE than LPR methods, particularly with small sample sizes.
- Substitution methods exhibited significant bias but sometimes low rMSE for small samples (<20).
- Non-parametric methods, including KM, performed poorly in contaminated distribution scenarios.
- Robust MLE and LPR versions showed less bias with contaminated data and multiple LODs.
Conclusions:
- Standard MLE demonstrated superior overall performance based on rMSE for both 95th percentile and mean estimation.
- Standard LPR outperformed robust LPR for mean estimation.
- Robust MLE methods are recommended when bias is the primary concern.
- The Kaplan-Meier method is not recommended for estimating the 95th percentile or mean in these scenarios.
Related Concept Videos
Censoring Survival Data
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different reasons...
Comparing the Survival Analysis of Two or More Groups
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and Cox...
Friedman Two-way Analysis of Variance by Ranks
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Survival Tree
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a survival tree begins...
Building a Survival Tree
Constructing a survival tree begins...
The Mantel-Cox Log-Rank Test
The Mantel-Cox log-rank test is a widely used statistical method for comparing the survival distributions of two groups. It tests whether a statistically significant difference exists in survival times between the groups without assuming a specific distribution for the survival data, making it a non-parametric test. This flexibility makes the log-rank test particularly valuable in medical research and other fields where the timing of an event, such as death or disease recurrence, is of interest.
