Back

A spatial EHR and wastewater-informed modeling framework for respiratory virus prediction under sparse and missing data conditions

Zhong, L.; Bleichrodt, A.; Pandey, A.; Kunkel, D.; Rennert, L.

2026-05-21 infectious diseases
10.64898/2026.05.18.26353485 medRxiv
Show abstract

Wastewater-based epidemiology has emerged as a powerful complement to clinical surveillance for monitoring infectious disease dynamics. However, most existing approaches either treat wastewater sites in isolation, overlooking spatial dependencies, and often fail to account for variability in data quality, limiting their ability to generate reliable predictions of healthcare demand. Here we present a spatial Bayesian renewal framework that integrates wastewater surveillance with mobility-informed spatial interactions while incorporating reliability-weighted wastewater signals. We apply the framework to three major respiratory pathogens, i.e., SARS-CoV-2, influenza, and respiratory syncytial virus (RSV), using wastewater and hospital data from counties in South Carolina. Across rolling four-week forecasts, the spatial framework consistently outperforms non-spatial approaches and remains robust even in counties lacking direct wastewater or hospitalization observations. Importantly, we show that county-level forecasts can be translated into facility-level predictions, enabling localized assessment of healthcare demand. These forecasts provide actionable early-warning signals to support hospital capacity planning, staffing decisions, and resource allocation. Together, this work establishes a scalable digital surveillance framework that integrates heterogeneous data sources for enabling more reliable infectious disease forecasting and supporting public health decision-making in underserved and data-limited settings.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.