Related Experiment Video
Updated: Jul 4, 2025

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
Better Estimates from Binned Income Data: Interpolated CDFs and Mean-Matching
Paul T von Hippel1, David J Hunter2, McKalie Drown2
1University of Texas at Austin.
Abstract:
Researchers often estimate income statistics from summaries that report the number of incomes in bins such as $0 to 10,000, $10,001 to 20,000, …, $200,000+. Some analysts assign incomes to bin midpoints, but this treats income as discrete. Other analysts fit a continuous parametric distribution, but the distribution may not fit well. We fit nonparametric continuous distributions that reproduce the bin counts perfectly by interpolating the cumulative distribution function (CDF). We also show how both midpoints and interpolated CDFs can be constrained to reproduce the mean of income when it is known. We evaluate the methods in estimating the Gini coefficients of all 3,221 U.S. counties. Fitting parametric distributions is very slow. Fitting interpolated CDFs is much faster and slightly more accurate. Both interpolated CDFs and midpoints give dramatically better estimates if constrained to match a known mean. We have implemented interpolated CDFs in the "binsmooth" package for R. We have implemented the midpoint method in the "rpme" command for Stata. Both implementations can be constrained to match a known mean.
Related Concept Videos
Measures of Central Tendency
Trimmed Mean
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Skewness
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
Distributions to Estimate Population Parameter
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Sampling Distribution

