Related Experiment Video
Updated: Oct 1, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
ASSESSING SELECTION BIAS IN REGRESSION COEFFICIENTS ESTIMATED FROM NONPROBABILITY SAMPLES WITH APPLICATIONS TO
Brady T West1, Roderick J Little2, Rebecca R Andridge3
1Survey Research Center, Institute for Social Research, University of Michigan.
New measures quantify selection bias in nonprobability samples, crucial for accurate genetic and survey research. These tools help assess how volunteer or user data might skew findings on polygenic scores and employment duration.
Area of Science:
- Statistics
- Biostatistics
- Social Sciences
Background:
- Selection bias poses a significant threat to the validity of inferences drawn from nonprobability samples.
- This bias is particularly relevant in genetic studies using volunteer samples for polygenic score (PGS) analysis and in surveys of specific user groups, like smartphone users.
Purpose of the Study:
- To derive novel measures for quantifying selection bias in regression models (linear and probit) fitted to nonprobability samples.
- To enable sensitivity analyses of inferences to assumptions about nonignorable selection when aggregate-level auxiliary data are available.
Main Methods:
- Development of selection bias measures derived from normal pattern-mixture models.
- Application of these measures to assess bias in polygenic score-phenotype relationships in a Facebook volunteer sample.
- Quantification of bias in subgroup mean differences (employment duration) within a smartphone user survey.
Main Results:
- The proposed measures effectively quantify selection bias in both simulated and real-world nonprobability samples.
- Bias was identified and quantified in estimated polygenic score-phenotype relationships and subgroup employment duration differences.
- Performance evaluation using benchmark estimates from large probability samples confirmed the measures' utility.
Conclusions:
- The novel measures provide a robust framework for assessing and understanding selection bias in nonprobability samples.
- These tools are essential for improving the reliability of findings in genetic epidemiology and survey research relying on convenience samples.
Related Concept Videos
Bias in Epidemiological Studies
Regression Toward the Mean
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Statistical Methods for Analyzing Epidemiological Data
Genetic Drift

