Related Experiment Video
Updated: Dec 28, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Bias due to Berkson error: issues when using predicted values in place of observed covariates
Gregory Haber1, Joshua Sampson1, Barry Graubard1
1Division of Cancer Epidemiology and Genetics, National Cancer Institute, 9609 Medical Center Drive, Bethesda, MD 20892, USA.
Abstract:
Studies often want to test for the association between an unmeasured covariate and an outcome. In the absence of a measurement, the study may substitute values generated from a prediction model. Justification for such methods can be found by noting that, with standard assumptions, this is equivalent to fitting a regression model for an outcome variable when at least one covariate is measured with Berkson error. Under this setting, it is known that consistent or nearly consistent inference can be obtained under many linear and nonlinear outcome models. In this article, we focus on the linear regression outcome model and show that this consistency property does not hold when there is unmeasured confounding in the outcome model, in which case the marginal inference based on a covariate measured with Berkson error differs from the same inference based on observed covariates. Since unmeasured confounding is ubiquitous in applications, this severely limits the practical use of such measurements, and, in particular, the substitution of predicted values for observed covariates. These issues are illustrated using data from the National Health and Nutrition Examination Survey to study the joint association of total percent body fat and body mass index with HbA1c. It is shown that using predicted total percent body fat in place of observed percent body fat yields inferences which often differ significantly, in some cases suggesting opposite relationships among covariates.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Random Error
Confounding in Epidemiological Studies
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Fundamental Attribution Error