Interpreting observational studies: why empirical calibration is needed to correct p-values
Martijn J Schuemie1, Patrick B Ryan, William DuMouchel
1Department of Medical Informatics, Erasmus University Medical Center Rotterdam, Rotterdam, The Netherlands; Observational Medical Outcomes Partnership, Foundation for the National Institutes of Health, Bethesda, MD, U.S.A.
Many observational studies incorrectly claim statistical significance due to bias. Empirical calibration of p-values is crucial for accurate drug safety assessments, revealing that over half of significant findings may be spurious.
Area of Science:
- Pharmacoepidemiology
- Biostatistics
- Medical Research Methodology
Background:
- The common reliance on p < 0.05 in medical literature assumes a 5% chance of false positives.
- Observational studies, unlike randomized trials, are particularly susceptible to bias and confounding, potentially invalidating this assumption.
- The integrity of statistical significance in drug safety research using observational data requires rigorous validation.
Purpose of the Study:
- To assess the validity of the p < 0.05 threshold in observational drug safety studies.
- To evaluate the impact of bias and confounding on statistical significance in real-world data.
- To introduce and test empirical calibration of p-values to account for systematic error.
Main Methods:
- Replication of three exemplar observational drug safety studies (case-control, cohort, self-controlled case series).
- Application of these designs to negative control drugs (drugs not expected to cause the outcome).
- Calculation of calibrated p-values incorporating random and systematic error, followed by literature analysis.
Main Results:
- A high frequency of spurious statistical significance (p < 0.05) was observed when no true effect was present.
- Empirical calibration effectively reduced false positive findings to the nominal 5% level.
- Analysis suggests at least 54% of literature findings with p < 0.05 may not be statistically significant.
Conclusions:
- The standard p < 0.05 threshold is unreliable in many observational drug safety studies.
- Empirical calibration is a vital tool for correcting systematic error and improving the reliability of research findings.
- A significant proportion of published observational study results warrant reevaluation due to potential false positives.
Related Concept Videos
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Bias in Epidemiological Studies
Statistical Significance
Calibration Curves: Correlation Coefficient
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...


