Related Experiment Video
Updated: Aug 17, 2025

Swabbing the Urban Environment - A Pipeline for Sampling and Detection of SARS-CoV-2 From Environmental Reservoirs
Published on: April 9, 2021
"Won't get fooled again": statistical fault detection in COVID-19 Latin American data
Dalson Figueiredo Filho1, Lucas Silva2, Hugo Medeiros1
1Department of Political Science, Universidade Federal de Pernambuco, Recife, Pernambuco, Brazil.
Background:
Claims of inconsistency in epidemiological data have emerged for both developed and developing countries during the COVID-19 pandemic.
Methods:
In this paper, we apply first-digit Newcomb-Benford Law (NBL) and Kullback-Leibler Divergence (KLD) to evaluate COVID-19 records reliability in all 20 Latin American countries. We replicate country-level aggregate information from Our World in Data.
Results:
We find that official reports do not follow NBL's theoretical expectations (n = 978; chi-square = 78.95; KS = 4.33, MD = 2.18; mantissa = .54; MAD = .02; DF = 12.75). KLD estimates indicate high divergence among countries, including some outliers.
Conclusions:
This paper provides evidence that recorded COVID-19 cases in Latin America do not conform overall to NBL, which is a useful tool for detecting data manipulation. Our study suggests that further investigations should be made into surveillance systems that exhibit higher deviation from the theoretical distribution and divergence from other similar countries.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Bias in Epidemiological Studies
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Detection of Gross Error: The Q Test
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

