Back

A Multi-pathogen Hospitalization Forecasting Model for the United States: An Optimized Geo-Hierarchical Ensemble Framework

Xu, S.; Du, H.; Dong, E.; Wang, X.; Zhang, L.; Gardner, L. M.

2025-01-08 infectious diseases
10.1101/2025.01.08.25320187 medRxiv
Show abstract

Accurate forecasting of infectious diseases is crucial for timely public health response. Ensemble frameworks have shown promising outcomes in short-term forecasting of COVID-19, among other respiratory viruses, however, there is a need to further improve these frameworks. Here, we propose the Multi-Pathogen Optimized Geo-Hierarchical Ensemble Framework (MPOG-Ensemble), a novel forecasting machine learning framework to forecast state-level hospitalizations of influenza, COVID-19, and RSV in the U.S. This framework is multi-resolution: it integrates state, regionally-trained, and nationally-trained models through an ensemble layer and applies various optimization methods to parameterize the model weights and enhance overall predictive accuracy. This proposed framework builds on existing forecasting literature by 1) employing an ensemble of three spatially hierarchical models with state-level forecasts as the output; 2) incorporating four distinct weight optimization methods to generate the ensemble; 3) utilizing clustering methods to dynamically identify multi-state regions as a function of short-term and long-term hospitalization trends for the regionally-trained model; and 4) providing a generalized multi-pathogen framework to forecast the expected near-term hospitalizations from Influenza, RSV and COVID-19. Results demonstrate MPOG-Ensemble is a robust framework with relatively high performance. Extensive experimentation using historical multi-pathogen data highlights the predictive power of our framework compared to existing ensemble approaches. Its robust performance underscores the frameworks effectiveness and potential for improving and broadening infectious disease forecasting.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.