Back

Comparing existing and novel methods for estimating etiology-specific diarrheal disease incidence in hybrid surveillance studies

Garcia Quesada, M.; Breskin, A.; Platts-Mills, J. A.; Benkeser, D.; Pavlinac, P. B.; Galagan, S. R.; Waller, L. A.; Lopman, B. A.; Rogawski McQuade, E. T.

2025-11-13 epidemiology
10.1101/2025.11.11.25339698 medRxiv
Show abstract

Accurate estimates of infectious disease incidence are critical for designing studies of public health interventions, including vaccines. Hybrid surveillance studies estimate incidence by enrolling cases in medical facilities, estimating population denominators in the community, and adjusting for healthcare seeking behaviors, which is necessary to minimize bias. The Enterics for Global Health (EFGH) Study aimed to generate updated incidence estimates of Shigella diarrhea among children in preparation for vaccine trials. We conducted a simulation to evaluate approaches for healthcare seeking adjustment and uncertainty estimation in hybrid studies and applied these methods to EFGH. Adjusting for healthcare seeking using the inverse of individual-level propensity scores for healthcare seeking greatly reduced bias compared to the inverse of the marginal probability for healthcare seeking. M-estimation and bootstrap 95% confidence intervals both had at least nominal coverage of the truth across scenarios. Monte Carlo 95% simulation intervals had nominal coverage in some scenarios but not all. When applied to EFGH, M-estimation confidence intervals around fully adjusted incidence estimates were narrower than bootstrap. Computation time for M-estimation using the geex R package was significantly higher than bootstrap or Monte Carlo, making bootstrap an appealing option for valid results and ease of use.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.