Related Experiment Video
Updated: Mar 7, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Performance of Firth-and logF-type penalized methods in risk prediction for small or sparse binary data
M Shafiqur Rahman1, Mahbuba Sultana2
1Institute of Statistical Research and Training, University of Dhaka, Dhaka, Bangladesh. shafiq@isrt.ac.bd.
For small or sparse datasets, penalized regression methods improve risk model performance over maximum likelihood estimation (MLE). The logF(1,1) method offers the best balance of calibration and discrimination for accurate risk prediction.
Area of Science:
- Statistics
- Biostatistics
- Machine Learning
Background:
- Maximum likelihood estimation (MLE) in logistic regression struggles with small/sparse data, leading to biased estimates and convergence failures due to separation.
- Separation issues, common even with large sample sizes and strong predictors, result in overfitted models with poor predictive performance.
- Firth- and logF-type penalized regression methods are alternatives to MLE for separation problems, but their use in risk prediction is limited.
Purpose of the Study:
- To evaluate Firth- and logF-type penalized regression methods for risk prediction in small or sparse datasets.
- To compare the performance of penalized methods against standard MLE and ridge regression.
- To assess calibration, discrimination, and overall predictive performance.
Main Methods:
- An extensive simulation study was conducted to evaluate predictive performance.
- Calibration, discrimination, and overall predictive performance were assessed for each method.
- A real-world dataset with low outcome prevalence was used for illustration.
Main Results:
- MLE demonstrated poor risk prediction performance in small or sparse datasets.
- All penalized methods showed improvements in calibration, discrimination, and overall performance compared to MLE.
- LogF(1,1) penalization outperformed other methods, including Firth-type and logF(2,2), with minimal bias and good predictive accuracy. Ridge regression showed good discrimination but often resulted in underfitted models and high convergence failure rates.
Conclusions:
- The logF-type penalized method, specifically logF(1,1), is recommended for practical use in developing risk models with small or sparse datasets.
- Penalized regression offers a viable solution to the challenges posed by separation in logistic regression for risk modeling.
- LogF(1,1) provides a robust approach for accurate risk prediction when data is limited or sparse.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Expected Frequencies in Goodness-of-Fit Tests
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Goodness-of-Fit Test

