The Effects of Data Preprocessing Choices on Behavioral RCT Outcomes: A Multiverse Analysis
1Center for Behavioural and Implementation Science, Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore.
Data preprocessing choices in randomized controlled trials (RCTs) significantly impact results, often more than statistical models. Transparent reporting and sensitivity analyses are crucial for robust behavioral science research.
Area of Science:
- Behavioral Science
- Biostatistics
- Data Science
Background:
- Data preprocessing decisions in randomized controlled trials (RCTs) can disproportionately influence study conclusions.
- Behavioral science data is often characterized by noise, skewness, and outliers, making preprocessing choices critical.
- The impact of preprocessing on RCT outcomes, especially in behavioral research, requires thorough investigation.
Purpose of the Study:
- To quantify the influence of data preprocessing pipelines on estimated treatment effects in simulated RCTs.
- To compare the impact of preprocessing choices versus model specification on study outcomes.
- To advocate for transparent reporting and sensitivity analyses of preprocessing steps in behavioral science RCTs.
Main Methods:
- Two multiverse analyses were conducted on simulated RCT data, encompassing 180 analytical pathways.
- Analyses crossed 36 preprocessing pipelines (varying outlier handling, imputation, and transformation) with five model specifications.
- Simulations utilized both linear regression families and advanced algorithms like generalized additive models, random forests, and gradient boosting.
Main Results:
- Preprocessing decisions explained a substantial majority of variance in estimated treatment effects (76.9% in linear models, 99.8% in advanced algorithms).
- Model specification had a minimal impact on variance (7.5% in linear models, 0.1% in advanced algorithms).
- Specific preprocessing pipelines drastically altered effect estimates, shrinking them by over 90% or inflating them by an order of magnitude.
Conclusions:
- Data preprocessing choices exert a far greater influence on RCT findings in behavioral science than statistical model selection.
- Meticulous reporting of preprocessing steps is essential for ensuring the robustness and replicability of research.
- Routine sensitivity or multiverse analyses are recommended to make the impact of preprocessing choices transparent.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
06:45Task Interruption and Resumption Paradigm for Testing the Activation and Pursuit of an Abstract Thinking Goal
Published on: April 18, 2017
Related Concept Videos
Regression Toward the Mean
Randomized Experiments
Simple randomization
Simple...
Group Design
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Blinding
