Related Experiment Video
Updated: Aug 30, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Measuring disparate outcomes of content recommendation algorithms with distributional inequality metrics
Tomo Lazovich1, Luca Belli1, Aaron Gonzales1
1Twitter, Inc., San Francisco, CA 94103, USA.
Abstract:
The harmful impacts of algorithmic decision systems have recently come into focus, with many examples of machine learning (ML) models amplifying societal biases. In this paper, we propose adapting income inequality metrics from economics to complement existing model-level fairness metrics, which focus on intergroup differences of model performance. In particular, we evaluate their ability to measure disparities between exposures that individuals receive in a production recommendation system, the Twitter algorithmic timeline. We define desirable criteria for metrics to be used in an operational setting by ML practitioners. We characterize engagements with content on Twitter using these metrics and use the results to evaluate the metrics with respect to our criteria. We also show that we can use these metrics to identify content suggestion algorithms that contribute more strongly to skewed outcomes between users. Overall, we conclude that these metrics can be a useful tool for auditing algorithms in production settings.
Related Concept Videos
Relative Frequency Distribution
Uniform Distribution
Two essential properties of this distribution are
Quantifying and Rejecting Outliers: The Grubbs Test
Skewness
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
Outliers and Influential Points
One-Way ANOVA: Unequal Sample Sizes

