Related Experiment Video
Updated: Dec 20, 2025

08:13
Using the Race Model Inequality to Quantify Behavioral Multisensory Integration Effects
Published on: May 10, 2019
6.7K
Differential privacy in the 2020 US census: what will it do? Quantifying the accuracy/privacy tradeoff
Samantha Petti1, Abraham Flaxman2
1School of Mathematics, Georgia Institute of Technology, Atlanta, GA, 30332, USA.
Gates Open Research
|June 2, 2020
Summary
The TopDown disclosure avoidance system for the US Census introduces minimal privacy loss, comparable to large random samples. This research aids in balancing data privacy and accuracy for future census data collection.
Area of Science:
- Data privacy and security
- Statistical methodology
- Demographic research
Background:
- The 2020 US Census utilizes the novel TopDown algorithm for disclosure avoidance.
- The TopDown algorithm was tested in the 2018 end-to-end census test.
- Census Bureau has publicly released the code and data from the 2018 test.
Purpose of the Study:
- To analyze the error introduced by the TopDown disclosure avoidance system.
- To empirically measure privacy loss in the TopDown approach.
- To compare TopDown's privacy and error with simple random sampling.
Main Methods:
- Utilized publicly available code and data from the 2018 end-to-end census test.
- Applied the TopDown algorithm to 1940 US census data.
- Developed an empirical measure of privacy loss for comparison.
Main Results:
- TopDown's empirical privacy loss is significantly lower than its theoretical guarantee.
- TopDown with a privacy budget of 1.0 showed error and privacy loss similar to a 50% simple random sample.
- TopDown with a privacy budget of 4.0 showed error and privacy loss similar to a 90% simple random sample.
Conclusions:
- The study contributes to the ongoing discussion on balancing privacy and accuracy in census data.
- Further research and discussion are needed to optimize data collection methods.
- The TopDown algorithm shows promise for effective disclosure avoidance with manageable error.
Related Concept Videos
Statistical Analysis: Overview
13.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
13.7K
Detection of Gross Error: The Q Test
6.7K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.7K
Uncertainty in Measurement: Accuracy and Precision
99.0K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.0K
Margin of Error
6.7K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
6.7K
Censoring Survival Data
442
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
442
Uncertainty: Confidence Intervals
9.5K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
9.5K

