Back

Learning shared forecast-error structure to improve ensemble forecasts of seasonal respiratory outbreaks

Qin, Y.; Du, H.; Pei, S.

2026-07-22 infectious diseases
10.64898/2026.07.20.26358508 medRxiv
Show abstract

Real-time forecasts of seasonal respiratory outbreaks are critical for public health preparedness and healthcare planning. Multi-model ensembles, which combine predictions from individual models, have become a leading approach for operational outbreak forecasting. Their success, however, depends in part on the assumption that component models make sufficiently independent errors. Here, we examined this assumption using archived real-time forecasts for influenza hospitalizations and influenza-like illness (ILI) in the United States. We found that component models with diverse structures and calibration methods shared systematic forecast errors during epidemic growth and around epidemic peaks, reflecting the common challenge of tracking rapid changes in epidemic dynamics from real-time surveillance data. Because such shared errors cannot be fully corrected by ensembling alone, we developed a deep learning framework that learns structured residual errors from historical forecasts and uses them to correct ensemble predictions. This framework improved influenza hospitalization forecasts across horizons and geographic scales, reducing the Weighted Interval Score by up to 20\% at the national level and 12\% across states relative to official ensemble forecasts, with the largest improvements at the near-term horizon and during epidemic growth and peak periods. We further showed that learned residual structures transferred across ensembles formed from different component models, making the approach robust to changes in model participation across seasons. The framework also improved ensemble forecasts for ILI, although gains were more modest. These findings reveal a fundamental challenge in ensemble forecasting and provide a generalizable approach for improving real-time epidemic forecasts.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.