Related Experiment Video
Updated: Sep 9, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Fully Randomized Predictor-Outcome Pairings Using a National Database Yield Frequent Statistical Significance Without
Whisper Grayson1, Aritra Chakraborty2, Nicholas M Brown1
1Department of Orthopaedic Surgery & Rehabilitation, Loyola University Health System, Maywood, Illinois.
Background:
Large national databases have enabled extensive outcomes research in arthroplasty. Their vast sample sizes, however, raise concerns regarding the identification of statistically significant, but clinically meaningless associations. We hypothesized that completely random pairings between database variables, without any clinical rationale, would still frequently yield statistically significant associations, which were driven purely by sample size rather than any meaningful relationship. This study examines the extent of false discovery risk in data-driven analysis of large clinical datasets.
Methods:
A retrospective cross-sectional analysis was performed utilizing a large national database. Patients who underwent total knee arthroplasty (TKA) or total hip arthroplasty (THA) were identified based on the Current Procedural Terminology (CPT) codes. There were 20 predictor-outcome variable pairs randomly selected using a seeded random generator applied across all variables in the dataset. Appropriate statistical tests were applied based on the variable types: Pearson correlations for continuous-continuous comparisons, Welch two-sample t-tests for continuous-binary comparisons, and Chi-square tests for binary-binary comparisons. Among 955,092 total cases (total hip arthroplasty: 373,528 [39.1%], total knee arthroplasty: 581,564 [60.9%]), mean age of 66 years (range, 17 to 89); mean body mass index (BMI) of 31.8 (range, 10.0 to 99.3); 58.6% of women, and 20 fully randomized predictor-outcome pairs were analyzed.
Results:
There were 14 variable pairs (70%) that yielded statistically significant results (P < 0.05), including Current Procedural Terminology code versus hemoglobin A1c (r = -0.0435, P = 2.14e-40), height versus history of chronic obstructive pulmonary disease (COPD) (change = 0.330, P = 1.83e-47), and case identification versus Current Procedural Terminology code (r = -0.027, P = 3.42e-150), despite no underlying clinical rationale.
Conclusions:
Fully randomized predictor-outcome pairings in a large national database yielded frequent statistically significant results without clinical relevance, which were driven purely by sample size. This highlights the risk of over-interpreting significance in large database studies and reinforces the need for hypothesis-driven methodology and effect size interpretation.
Related Concept Videos
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Statistical Significance
Cochran's Q Test
McNemar's Test
Wilcoxon Signed-Ranks Test for Matched Pairs
Significance Testing: Overview

