Back

Robust uncertainty quantification in popular estimators of the instantaneous reproduction number

Steyn, N.; Parag, K. V.

2024-10-22 infectious diseases
10.1101/2024.10.22.24315918 medRxiv
Show abstract

The instantaneous reproduction number (Rt) is a key measure of the rate of spread of an infectious disease. Correctly quantifying uncertainty in Rt estimates is crucial for making well-informed decisions. Popular Rt estimators leverage smoothing techniques to distinguish signal from noise. Examples include EpiEstim and EpiFilter, which are both controlled by a "smoothing parameter" that is traditionally selected by users. We demonstrate that the values of these smoothing parameters are unknown, vary markedly with epidemic dynamics, and show that data-driven smoothing is crucial for accurate uncertainty quantification of Rt estimates. We derive model likelihoods for the smoothing parameters in both EpiEstim and EpiFilter and develop a Bayesian framework to automatically marginalise these parameters when fitting to epidemiological time-series data. This yields novel marginal posterior predictive distributions which prove integral to rigorous model evaluation. Applying our methods, we find that default parameterisations of these widely-used estimators can negatively impact Rt inference, delaying detection of epidemic growth, and misrepresenting uncertainty (typically producing overconfident estimates), with implications for public health decision-making. Our extensions mitigate these issues, provide a principled approach to uncertainty quantification, improve the robustness of real-time Rt inference, and facilitate model comparison using observable quantities.

Published in American Journal of Epidemiology (predicted rank #17) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.