Related Experiment Videos
Reproducibility of social science research using aggregate statistics with noise infused for differential privacy
Ryan Steed1, Eduardo Abraham Schnadower Mustri1, Alessandro Acquisti2
1Heinz College of Information Systems and Public Policy, Carnegie Mellon University, Pittsburgh, PA 15213.
Abstract:
Privacy-preserving analytics such as differential privacy are designed to allow the analysis of sensitive datasets while protecting individuals' privacy. Their deployment, however, has been controversial. Critics maintain that statistical noise injected to preserve privacy can degrade the quality and feasibility of social science research. We select a benchmark of empirical findings from 93 published social science studies involving regression analyses over aggregate statistics. We evaluate whether their findings replicate on privacy noise-infused data. Under privacy budgets typical in industry, around 91% of simulated findings still support the original claims at significance level [Formula: see text]. Claims based on weaker original effect sizes are more likely to be nullified or sometimes reversed. We compare distortions caused by privacy noise to those due to measurement errors and other kinds of nonsampling errors common in social statistics, and we find that the marginal impacts of privacy protection are smaller. Moreover, discrepancies due to privacy noise are often much smaller than discrepancies observed in traditional replication and robustness studies.
Related Concept Videos
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Statistical Methods for Analyzing Epidemiological Data
Regression Toward the Mean
Testing a Claim about Mean: Known Population SD
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Random Error
Empirical Method to Interpret Standard Deviation
This rule is used widely in statistics to calculate the proportion of data values...