Quantifying the information in noisy epidemic curves
Kris V Parag1,2, Christl A Donnelly3,4, Alexander E Zarebski5
1NIHR Health Protection Research Unit in Behavioural Science and Evaluation, University of Bristol, Bristol, UK. kris.parag@bristol.ac.uk.
Abstract:
Reliably estimating the dynamics of transmissible diseases from noisy surveillance data is an enduring problem in modern epidemiology. Key parameters are often inferred from incident time series, with the aim of informing policy-makers on the growth rate of outbreaks or testing hypotheses about the effectiveness of public health interventions. However, the reliability of these inferences depends critically on reporting errors and latencies innate to the time series. Here, we develop an analytical framework to quantify the uncertainty induced by under-reporting and delays in reporting infections, as well as a metric for ranking surveillance data informativeness. We apply this metric to two primary data sources for inferring the instantaneous reproduction number: epidemic case and death curves. We find that the assumption of death curves as more reliable, commonly made for acute infectious diseases such as COVID-19 and influenza, is not obvious and possibly untrue in many settings. Our framework clarifies and quantifies how actionable information about pathogen transmissibility is lost due to surveillance limitations.
Related Concept Videos
Steps in Outbreak Investigation
Statistical Methods for Analyzing Epidemiological Data
Causality in Epidemiology
Survival Curves
The Kaplan-Meier estimator is the most common method for constructing survival curves. This...
Censoring Survival Data
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...


