Related Experiment Videos
Estimation of probabilities using the logistic model in retrospective studies
1School of Mathematics and Science, California State University, Northridge 91330.
Computers and Biomedical Research, an International Journal
|October 1, 1988
Summary
Four methods for estimating logistic regression constant terms in case-control studies were compared. Method 1, maximum likelihood estimation, best estimated posterior probabilities and improved classification accuracy for binary variables.
Area of Science:
- Biostatistics
- Statistical Modeling
Background:
- Case-control studies are common in epidemiology and clinical research.
- Logistic regression is a widely used statistical model for binary outcomes.
- Estimating parameters in logistic regression with case-control data presents unique challenges, particularly for the constant term.
Purpose of the Study:
- To compare four methods for estimating the constant term in logistic regression models using case-control data.
- To evaluate the performance of these estimators for both large and small sample sizes.
- To assess the impact of different constant term estimation methods on the accuracy of posterior probability (Px) estimation and classification procedures.
Main Methods:
- Maximum likelihood estimation (MLE) was used for regression coefficients.
- Four distinct methods were proposed and compared for estimating the constant term parameter.
- Asymptotic distribution analysis was employed for large sample comparisons.
- Simulation studies using 11 logistic models were conducted for small sample comparisons and classification performance evaluation.
Main Results:
- Method 1 (MLE-based) demonstrated the minimum expected mean square error when estimating the posterior probability (Px).
- Performance differences were less clear when the constant term itself was the primary parameter of interest.
- For logistic models with predominantly binary predictors, Method 1 yielded the minimum expected error rate in classification.
- In other cases, the linear discriminant function (LDF) provided the minimum expected error rate, with the logistic discriminant using Method 1 being comparable.
Conclusions:
- Method 1 is recommended for estimating posterior probabilities in logistic regression with case-control data.
- The choice of method for estimating the constant term impacts classification accuracy, particularly concerning the nature of predictor variables.
- The LDF may be preferable when predictors are not predominantly binary, though logistic discriminant analysis with Method 1 remains a strong alternative.