Related Experiment Video
Updated: Jul 4, 2025

12:39
A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
11.4K
Better Estimates from Binned Income Data: Interpolated CDFs and Mean-Matching
Paul T von Hippel1, David J Hunter2, McKalie Drown2
1University of Texas at Austin.
Summary
This study introduces a faster, more accurate method for estimating income distributions from binned data using interpolated cumulative distribution functions (CDFs). Constraining these estimates to a known mean significantly improves accuracy for income statistics like Gini coefficients.
Area of Science:
- Economics
- Statistics
- Data Science
Background:
- Estimating income statistics from binned data is common.
- Existing methods like bin midpoints or parametric distributions have limitations in accuracy and speed.
- Accurate income distribution estimation is crucial for socioeconomic analysis.
Purpose of the Study:
- To develop and evaluate improved methods for estimating income statistics from binned data.
- To compare the performance of nonparametric interpolated cumulative distribution functions (CDFs) against traditional methods.
- To assess the impact of constraining estimates to a known mean on accuracy.
Main Methods:
- Fitting nonparametric continuous distributions by interpolating the cumulative distribution function (CDF) to match bin counts.
- Constraining both interpolated CDFs and bin midpoints to reproduce a known mean income.
- Evaluating Gini coefficient estimation accuracy across 3,221 U.S. counties.
Main Results:
- Interpolated CDFs accurately reproduce bin counts and are faster than parametric methods.
- Constraining estimates to a known mean dramatically improves accuracy for both interpolated CDFs and midpoints.
- Interpolated CDFs offer a slight accuracy improvement over constrained midpoints.
Conclusions:
- Nonparametric interpolated CDFs provide a superior method for estimating income distributions from binned data.
- Matching estimates to a known mean is a critical step for enhancing the reliability of income statistics.
- Software packages 'binsmooth' (R) and 'rpme' (Stata) are available for implementing these methods.
Related Concept Videos
Measures of Central Tendency
16.0K
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians,...
16.0K
Trimmed Mean
2.9K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
2.9K
Skewness
11.1K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
11.1K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Weighted Mean
5.2K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.2K
Sampling Distribution
12.6K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
12.6K

