Back

Forecastability of infectious disease time series: are some seasons and pathogens intrinsically more difficult to forecast?

White, L. A.; Leon, T. M.

2025-04-30 public and global health
10.1101/2025.04.29.25326677 medRxiv
Show abstract

For infectious disease forecasting challenges, individual model performance typically varies across space and time. This phenomenon raises the question: are there properties of the target time series that contribute to a particular season, location, or disease being more difficult to forecast? Here we characterize a time series future predictability using a forecastability metric that calculates the spectral density of the time series. Forecastability of syndromic influenza hospital admissions for the state of California varied widely across seasons and was positively correlated with peak burden. Next, using archived U.S. state and national forecasts targeting laboratory-confirmed COVID-19 and influenza hospital admissions, we investigated the relationship between forecastability and: (i) population size of the forecasting target, and (ii) forecast performance as measured by mean absolute error, weighted interval score (WIS), and scaled relative WIS. Forecastability increased with increasing population size of the forecasting target, and forecasting performance generally improved with higher forecastability when controlling for population size across scales. These preliminary results support the idea that some targets and respiratory virus seasons may be inherently more difficult to forecast and could help explain inter-seasonal variation in model performance. Author summaryCould intrinsic properties of an epidemiological time series help explain why a particular season, location, or disease is more difficult to predict in the future? To answer this question, this analysis uses a measure of a time series future predictability called "forecastability," which describes the inherent uncertainty or surprise in the signal. Influenza and COVID-19 hospital admissions had higher forecastability scores for locations with larger population sizes, possibly due to larger counts leading to smoother time series. At the same time, forecasting performance generally improved for time series with higher forecastability scores when controlling for population size, suggesting that this metric is helpful for understanding ease of forecasting. These preliminary results support the idea that some epidemiological targets and respiratory virus seasons may be inherently more difficult to forecast and could help explain why forecasting model performance changes across different respiratory virus seasons.

Published in PLOS Computational Biology (predicted rank #3) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.