Related Experiment Video
Updated: May 29, 2026

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
Published on: January 13, 2023
An automatic finite-sample robustness metric: when can dropping a little data change conclusions? Part I: definitions
Ryan Giordano1, Rachael Meager2, Tamara Broderick3
1Department of Statistics, University of California, Berkeley, CA, USA.
None:
Study samples often differ in non-random ways from the target populations to which policy decisions will eventually be applied. Researchers typically hope that such departures from random sampling-due to changes in the population over time and space, or difficulties in sampling truly randomly-are small, and their corresponding impact on the inference should be small as well. Accordingly, researchers might be concerned if the conclusions of their studies are excessively sensitive to a very small proportion of our sample data. We propose a method to assess the sensitivity of applied conclusions to the removal of a small fraction of the sample. Manually checking the influence of all possible small subsets is computationally infeasible, so we use an approximation to find the most influential subset. Our metric, the 'Approximate Maximum Influence Perturbation', is based on the classical influence function. It is automatically computable for common methods including (but not limited to) ordinary least squares, instrumental variables regression, maximum likelihood, generalized method of moments and variational Bayes. At minimal extra cost, we provide an exact finite-sample lower bound on sensitivity. While some empirical applications are robust, we show that results of several influential economics papers can be overturned by removing less than 1% of the sample. This article is part of the theme issue 'Statistical workflow'.
Related Concept Videos
Censoring Survival Data
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Detection of Gross Error: The Q Test
Regression Toward the Mean
Quantifying and Rejecting Outliers: The Grubbs Test
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...