Related Experiment Video
Updated: Sep 9, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Fully Randomized Predictor-Outcome Pairings Using a National Database Yield Frequent Statistical Significance Without
Whisper Grayson1, Aritra Chakraborty2, Nicholas M Brown1
1Department of Orthopaedic Surgery & Rehabilitation, Loyola University Health System, Maywood, Illinois.
Large national databases can produce statistically significant, but clinically meaningless, results due to sheer sample size. This study shows that random variable pairings frequently yield false discoveries, highlighting the need for hypothesis-driven research.
Area of Science:
- Orthopedic Surgery
- Health Informatics
- Biostatistics
Background:
- Large national databases are valuable for arthroplasty outcomes research.
- Vast sample sizes in these databases can lead to statistically significant but clinically irrelevant findings.
- This study investigates the risk of false discoveries in data-driven analyses of large clinical datasets.
Purpose of the Study:
- To test the hypothesis that random variable pairings in large datasets yield statistically significant associations.
- To examine the extent of false discovery risk in data-driven analysis of large clinical datasets.
Main Methods:
- Retrospective cross-sectional analysis of a large national database.
- Identification of patients undergoing total knee (TKA) or total hip arthroplasty (THA) using Current Procedural Terminology (CPT) codes.
- Random selection of 20 predictor-outcome variable pairs for analysis, with appropriate statistical tests applied based on variable types.
Main Results:
- 70% (14 out of 20) of randomly paired variables yielded statistically significant results (P < 0.05).
- Examples include CPT code vs. hemoglobin A1c and height vs. chronic obstructive pulmonary disease (COPD).
- These significant associations were found despite a lack of underlying clinical rationale.
Conclusions:
- Random pairings in large databases frequently produce statistically significant results without clinical relevance, driven by sample size.
- This underscores the risk of over-interpreting statistical significance in large database studies.
- Emphasizes the necessity of hypothesis-driven methodology and effect size interpretation in outcomes research.
Related Concept Videos
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Statistical Significance
Cochran's Q Test
McNemar's Test
Wilcoxon Signed-Ranks Test for Matched Pairs
Significance Testing: Overview

