从分类收入数据中获得更好的估计:互插的CDF和平均匹配
Paul T von Hippel1, David J Hunter2, McKalie Drown2
1University of Texas at Austin.
概括
这项研究引入了一种更快,更准确的方法,用于估计使用插入累积分布函数 (CDFs) 来估计来自内存数据的收入分配. 将这些估计限制在已知的平均值上,可以显著提高收入统计的准确性,例如吉尼系数.
科学领域:
- 经济学 经济学 经济学
- 统计 统计 统计 统计
- 数据科学数据科学数据科学
背景情况:
- 估计收入统计数据从内存数据是常见的.
- 像 bin 中点或参数分布这样的现有方法在准确性和速度上有局限性.
- 准确的收入分配估计对于社会经济分析至关重要.
研究的目的:
- 开发和评估改进的方法来估计收入统计数据.
- 为了比较非参数交叉累积分布函数 (CDF) 与传统方法的性能.
- 评估将估计限制在已知的平均值对准确性的影响.
主要方法:
- 通过插入累积分布函数 (CDF) 来匹配垃圾箱数量来适应非参数连续分布.
- 限制插入的CDF和 bin中点来复制已知的平均收入.
- 评估美国3221个县的吉尼系数估计准确度.
主要成果:
- 互波式CDF准确地复制了垃圾箱计数,并且比参数方法更快.
- 将估计限制在已知的平均值上大大提高了对插曲的CDF和中点的准确性.
- 互波式CDF比受约束的中点提供了轻微的精度改进.
结论:
- 非参数互波式CDF提供了一种优越的方法,用于估计来自内存数据的收入分配.
- 将估计与已知的平均值相匹配是提高收入统计数据可靠性的关键步骤.
- 软件包"binsmooth" (R) 和"rpme" (Stata) 可用于实施这些方法.
相关概念视频
Measures of Central Tendency
16.0K
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians,...
16.0K
Trimmed Mean
2.9K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
2.9K
Skewness
11.1K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
11.1K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Weighted Mean
5.2K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.2K
Sampling Distribution
12.6K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
12.6K


