关于"统计不确定性和隐私对政策的影响"的技术评论
Yifan Cui1, Ruobin Gong2, Jan Hannig3
1Center for Data Science, Zhejiang University, China.
概括
官方统计数据的质量对于政策决策至关重要. 这项研究批评了现有的评估方法,并提出了基于模拟的改进技术,以更可靠地评估数据质量.
科学领域:
- 统计数据
- 数据科学
- 公共政策
背景情况:
- 官方统计数据产品在准确性,稳定性和公平性方面显著影响政策决策.
- 即使精心整理的数据也可能包含错误或不准确.
- 统计数据的质量直接影响了基于证据的决策可靠性.
研究的目的:
- 强调对官方统计数据产品进行原则性质量评估的必要性.
- 确定Steed等人使用的质量评估方法的局限性.
- 提出和讨论数据质量评估的其他统计学合理方法.
主要方法:
- 对估计者的可接受性和Steed等人中的诱导概率模型的批评. 这是一个很好的评估.
- 开发基于模拟的方法来估计允许的最小收缩.
- 应用多层实证贝叶斯模型进行质量评估.
主要成果:
- 通过Steed等进行的质量评估程序. 显示出统计不可接受性和模型不一致性.
- 建议的替代方法为评估统计数据质量提供了更有原则的方法.
- 基于模拟的技术在评估数据产品方面显示出更高的可靠性.
结论:
- 正确评估官方统计数据的质量对于明智的政策决策至关重要.
- 质量评估必须考虑到固有的数据不确定性和特定的下游使用情况.
- 需要改进的统计方法来确保官方数据的可靠性.
相关概念视频
Propagation of Uncertainty from Random Error
745
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
745
Statistical Analysis: Overview
6.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.7K
Statgraphics
168
Statgraphics is a comprehensive statistical software suite designed for both basic and advanced data analysis. Originating in 1980 at Princeton University under Dr. Neil W. Polhemus, it was one of the pioneering tools for statistical computing on personal computers, with its public release in 1982 marking an early milestone in data science software. Over the years, it has evolved into a robust platform for data science, offering tools for regression analysis, ANOVA, multivariate statistics,...
168
Random Error
937
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
937
Biostatistics: Overview
290
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
290
Uncertainty: Overview
607
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
607


