Epidemics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Epidemics's content profile, based on 116 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.
Parpia, A.; Wright, J.; Gharouni, A.; Thampi, N.; Fitzpatrick, T.
Show abstract
Background: Respiratory syncytial virus (RSV) remains a leading cause of hospitalization in infancy, with severe outcomes influenced by both contact patterns and passive immunity. Non-pharmaceutical interventions (NPIs) during the COVID-19 pandemic suppressed RSV circulation and reduced opportunities for maternal immune boosting, potentially altering protection among newborns. We evaluated whether incorporating time-varying maternal immunity improves the ability of an age-structured transmission model to predict post-pandemic RSV hospitalization patterns in infants. Methods: We analyzed population-based RSV hospitalizations among Ontario (Canada) infants (<1 year) from July 2, 2017 to June 25, 2024, using linked administrative databases. A deterministic compartmental model across seven age classes was calibrated against pre-pandemic data using Latin Hypercube Sampling. We compared a model incorporating time-varying contact rates alone against a specification that additionally included time-varying maternal immunity. Results: Both specifications accurately reproduced pre-pandemic seasonality and macro-level post-pandemic resurgence features. The constant maternal immunity model showed slightly better accuracy in capturing the 2021/22 peak compared to the time-varying maternal immunity specification. However, both qualitatively captured the continued near-absence of RSV and the observed peak was captured within the 95% credible intervals. While both models precisely captured the timing and overwhelming surge of admissions that occurred in 2022/23, they failed to capture the premature peak timing and magnitude in 2023/24. Conclusions: Incorporating time-varying maternal immunity did not improve model accuracy post-pandemic. While maternal protection is essential for evaluating infant immunizations, population-level contact shifts primarily shaped post-pandemic RSV seasonality, indicating that models must account for these mechanisms of RSV transmission dynamics.
Bastard, J.; Assaad, C.; Marti, R.; Tran, A.; Metras, R.; DURAND, B.
Show abstract
Models that provide risk maps for zoonoses often lack (i) a spatiotemporal autocorrelation component, yet crucial in understanding the spread of infectious diseases, (ii) accounting for heterogeneity in case reporting, and (iii) a causal framework for explanatory variables. Here, we addressed these limitations with a model system, West Nile virus, a vector-borne pathogen transmitted in a bird reservoir, and affecting humans and horses. We built a spatiotemporal occupancy model and fitted it to notified (human and horse) case data. Based on a directed acyclic graph, we estimated the causal effects of conjectural weather variables (i.e. changing in the short-term) vs. structural variables (i.e. changing in the long-term) on WNV circulation in the bird reservoir, besides assessing variables associated with case reporting. By computing population attributable fractions, we found the contribution of conjectural weather variables to WNV outbreaks in Europe to be globally higher than the structure of the bird community.
Loo, S. L.; Nande, A.; Hill, A. L.; Truelove, S.
Show abstract
Age is a primary determinant of symptom severity and transmission patterns for many infectious diseases, motivating the use of age-stratified models parameterized by contact matrices. In the United States, the absence of direct contact surveys has required estimating synthetic contact matrices from demographic data on household size, school attendance, and workforce participation. However, this likely underestimates contacts among children under age 5, who often attend group childcare missing from censuses. The goal of this study was to use nationally-representative data on childcare arrangements (the Early Childhood Program Participation Survey) to reconstruct daily contacts occurring in childcare settings, and augment existing all-age contact matrices. For infants under 1 year of age, we estimated 0.2 daily contacts with other infants, increasing to 0.7 daily contacts with same-age peers for 1- or 2-year-olds, 1.3 for 3-year-olds, and 3.5 for 4-year-olds. Including childcare settings increases estimated contacts among young children by up to six fold. Using simulations of measles outbreaks in inadequately vaccinated populations, we show that prior contact matrices significantly underestimated the outbreak frequency, size, and impact on preschool age groups. Our findings highlight the need for targeted data collection on childcare contacts to improve model-based evaluation of interventions particularly for young children.
Pi, L.; Davis, E. L.; Danon, L.; Hollingsworth, D.
Show abstract
Long-term care facilities (LTCs) worldwide experienced disproportionately high infection and mortality rates during the COVID-19 pandemic, where essential care limits opportunities for contact segregation. However, empirical contact data remain scarce, limiting our understanding of how individual contact behaviours shape transmission in these settings. In this study, we developed a stochastic network-based transmission model parameterised using real-world self-reported contact data collected from a median-sized UK LTC unit. By incorporating high-resolution observational data that reflect routine care delivery patterns, we quantified how heterogeneity in contact networks influences outbreak dynamics. We found substantial variation in contact behaviour between individuals, resulting in highly heterogeneous transmission outcomes. Outbreak occurrence, timing, final size, and the likelihood of super-spreading events all varied markedly depending on the structure of the underlying contact network and the characteristics of the index case. Individuals with high contact activity were considerably more likely to initiate large outbreaks than those with fewer contacts. For a per-contact transmission probability of 10%, introduction of infection through the most highly connected individuals resulted in a greater than 75% probability of a large outbreak. Our findings indicate that preventing infection introduction through both residents and staff is critical for outbreak control in LTCs. Individuals with high contact activity were consistently associated with a greater probability of initiating large outbreaks, highlighting the importance of accounting for contact heterogeneity when designing surveillance and infection-control measures. More broadly, this study demonstrates the importance of accounting for contact-network heterogeneity when designing infection prevention and control measures in LTC settings, and highlights the value of integrating empirical contact data with transmission modelling to inform evidence-based outbreak preparedness, targeted surveillance, and infection-control strategies in long-term care facilities.
Sevilla, J.; Kende, J.; Duchene, S.; Meehan, M. T.
Show abstract
Bacterial sexually transmitted infections (STIs) pose a major global public health challenge, with Neisseria gonorrhoeae being of particular concern due to its persistently high prevalence and increasing antimicrobial resistance. The emergence of multidrug-resistant strains has narrowed treatment options, highlighting the importance of prevention. In this context, knowing whether there is superspreading (transmission heterogeneity) within a population becomes crucial for accurate public health measures. However, classic methods to quantify superspreading rely on dense contact tracing, and this is not always feasible. As an alternative, we can use Bayesian phylodynamic modelling to infer transmission dynamics, including superspreading. Yet modelling transmission dynamics using bacterial data remains problematic, although it is widely used for viral data. Here, we apply a multi-type birth-death model parametrised to quantify superspreading in N. gonorrhoeae outbreaks, estimating the fraction and relative impact of superspreaders and reproductive numbers for superspreaders and non-superspreaders. We also use a hierarchical modelling strategy with partial pooling to increase the power for detecting superspreading in each cluster. Model performance was successfully evaluated across a range of superspreading scenarios using both transmission-informed phylogenies and sequence data with phylogenetic uncertainty. Application to empirical genomic data revealed a substantial role of superspreading in N. gonorrhoeae transmission during the COVID-19 pandemic in Australia. These results highlight the impact of superspreading in N. gonorrhoeae transmission and the importance of detecting it to efficiently stop the dissemination of the disease
Wakwella, A. R.; Klein, C. J.; Woodberry, O.; Lau, C. L.; Wenger, A.; Jupiter, S. D.; Jenkins, A. P.; Mayfield, H. J.
Show abstract
Leptospirosis is a water-related zoonotic disease with complex transmission pathways, including direct transmission from infected animals and indirect transmission through contaminated soil and water. Identifying key areas to implement targeted infection prevention and control strategies is challenging, as a range of risk factors across different scales can drive human infection. We aimed to develop an epidemiological modelling approach to predict key transmission pathways driving leptospirosis infection, including risks ranging from household level factors to the movement of pathogens across watersheds. We combined a causal Bayesian network with a novel hydrological pathogen transport model to predict leptospirosis across Fiji and found key infection hotspots adjacent to rivers and within degraded watersheds; a dynamic overlooked by previous epidemiological models. We used a wide range of data for model parameterisation (e.g., expert elicitation, epidemiological surveys, and environmental data) and found that predictive validity improved when expert input on risk factors was used to guide model parameterisation, improving R2 for predicted versus observed seroprevalence from 0.71 to 0.91. Our One Health modelling approach can be used to support the design and evaluation of environment-based disease prevention strategies at national scales.
Verheyden, J. G. L.; Mudogo, C. N.; Jacquet, W.
Show abstract
Background: Real-time outbreak analyses are often requested before surveillance systems have stabilised or epidemics have generated enough information for the desired inference. Existing approaches address surveillance quality, forecasting, estimands and identifiability separately, but do not provide a common rule for deciding which analytical product is supportable at a particular data vintage. We developed an estimand-first framework for analytical readiness. Methods: The framework distinguishes surveillance maturity (S), epidemic-process informativeness (E) and estimand-specific analytical readiness, defined as whether the available data vintage, observation process, method and decision-matched validation support a specified inference for a specified decision. We stress-tested four implications using longitudinal data from the 2018-2020 Ebola response in eastern Democratic Republic of the Congo (DRC), archived geographic forecasts, independent forecasting data from Western Area, Sierra Leone, and a targeted mortality-identifiability experiment. Results: During a documented DRC surveillance disruption and recovery, seven-day persistence forecasts had all-health-zone absolute errors of 1, 3, 18 and 4 cases across pre-shock, acute-shock, early-recovery and recovery origins; the largest error occurred during early recovery. Four-week reported-case trend multipliers changed from 0.67 and 0.73 to 1.12 and 1.29, while the final fit was strongly overdispersed (Pearson dispersion 7.65), demonstrating asynchronous readiness across estimands. Archived geographic forecasts improved a Top-3 allocation decision over cumulative burden at only one origin despite consistently lower Brier scores for one specification. In Western Area, persistence forecast mean absolute error increased from 39.1 cases at one week to 82.3 at four weeks, and a history-to-horizon ratio did not define a universal threshold. An observed reported case-fatality ratio of 0.40 was compatible with constructed latent fatality values from 0.10 to 0.80; increasing the reported denominator narrowed sampling uncertainty without reducing structural uncertainty. Conclusions: Analytical readiness is task- and vintage-specific rather than a property of a dataset. More data, model convergence or narrow intervals cannot substitute for estimand definition, observation-process awareness, decision-matched validation and explicit identification analysis. Keywords: outbreak analytics; surveillance maturity; analytical readiness; estimand; identifiability; forecasting; Ebola; reporting process; decision-matched validation
Topazian, H. M.; Sheets, T. R.; Gruninger, R. J.; Kelley, J.; LaCross, N.; Samore, M. H.; Lofgren, E.; Keegan, L. T.
Show abstract
Since the COVID-19 pandemic, forecasting hubs and non-traditional respiratory disease surveillance streams have become increasingly common. However, many forecasting approaches assume that relationships between surveillance predictors and disease outcomes remain stable over time and that incorporating additional historical data will improve forecast performance. To evaluate these assumptions in a real-world setting, we developed and evaluated forecasts of SARS-CoV-2 and influenza hospitalizations in Utah using syndromic surveillance, test positivity, and wastewater data. Rather than identifying a single, best-performing model, we examined whether relationships between surveillance predictors and hospitalization outcomes remained stable across seasons and whether longer historical training periods consistently improved forecast accuracy. Relationships between surveillance predictors and hospitalizations varied substantially by pathogen and season. Analyses using pooled data across multiple years suggested strong positive correlations between predictors and outcomes, but these aggregated patterns often obscured weak or negative correlations observed during SARS-CoV-2 variant waves and influenza seasons. Forecast performance similarly varied over time. Models that performed well during some seasons, transmission phases, or under certain training strategies frequently performed worse than benchmark models in others. Training on additional historical data generally reduced forecast accuracy, though this varied by disease and transmission phase. Forecasting groups should prioritize continual evaluation of surveillance predictors, adaptive strategies, and diverse ensembles, rather than relying on a single model, data stream, or historical training framework each year.
Jamieson, N.; Charalambous, C.; Schultz, D. M.; Naik, F.; Howett, D.; Dabrera, G.; Hall, I.
Show abstract
Legionnaires disease is a severe respiratory illness caused by Legionella bacteria, with most cases occurring sporadically and environmental sources often unidentified. Effective outbreak detection requires understanding the spatiotemporal dynamics of sporadic cases and their environmental drivers. We developed a mechanistically informed spatiotemporal model integrating fine-scale spatial heterogeneity, multi-week meteorological influences, and extended temporal lags. The framework combines a negative binomial generalised additive model (GAM), a Besag-York-Mollie (BYM2) spatial component, and distributed lag nonlinear models (DLNMs) to capture nonlinear, delayed effects of temperature, dewpoint depression, precipitation, and cloud cover. These outputs generate a national daily index of weather-driven vulnerability, which is combined with hierarchical clustering to identify potential outbreaks. Across 2000-2019, our model improved outbreak detection in 15 of 20 years compared with the baseline UKHSA approach; in the remaining years performance was either equivalent (two years) or only slightly worse (three years, 0.98% reduction). Mean relative improvements were 6.0%, with a maximum of 14.4% in 2013. Improvements were consistent across months, and coarser 0.25-degree grid evaluations likely underestimate the models advantage at finer spatial scales. The analysis also clarified dual-stage Legionnaires disease dynamics, distinguishing environmental bacterial growth from the shorter infection window, and demonstrated the necessity of extended lags for accurate risk prediction. This framework provides a robust platform for targeted surveillance and predictive modelling, supporting evidence-based interventions and enhancing preparedness for sporadic Legionnaires disease under observed climatic conditions.
Verheyden, J. G. L.; Mudogo, C. N.
Show abstract
National-level estimates of the time-varying reproduction number (Rt) for the 2026 Bundibugyo virus disease (BDBV) outbreak in the Democratic Republic of the Congo (DRC) have converged on a value close to the epidemic threshold since early August 2026, consistent with independent joint Bayesian renewal-model estimates. A single national Rt, however, can obscure divergent sub-national epidemic trajectories, particularly across a five-province outbreak in which provinces range from a declining original epicentre to recently-seeded fronts. We estimated Rt at national, provincial, and, where case volume allowed, health-zone level, using both a standard sliding-window (Cori) estimator and a hierarchical Bayesian renewal model with partial pooling across spatial units, fitted by Hamiltonian Monte Carlo (No-U-Turn Sampler). Provincial estimates diverged materially from the national trend: as of the week of 6 August 2026, Ituri, the outbreak's original epicentre, had a hierarchical median Rt of 0.91 (95% credible interval [CrI] 0.68 - 1.24), while Nord-Kivu (1.23, [0.87 - 1.72]) and Haut-Uele (1.79, [1.24 - 2.67]) remained above threshold. Health-zone disaggregation, feasible only in Ituri and Nord-Kivu given case volume, showed this provincial picture itself masked further heterogeneity: in Ituri, the zone where the outbreak began (Mongbwalu) had clearly declined (Rt 0.36, [0.17 - 0.74]) while the two largest zones by cumulative case count (Bunia, Rwampara) remained at or above threshold; in Nord-Kivu, elevated transmission was concentrated in a single zone (Katwa, Rt 1.39, [0.85 - 1.99]) while a comparably-sized zone (Butembo) had already declined (0.64, [0.23 - 1.38]). An initial disagreement between the sliding-window and hierarchical provincial estimates was traced to a data-reconstruction artefact (forward-filling, rather than interpolating, multi-day gaps in health-zone reporting) rather than a genuine methods disagreement, and resolved once corrected. The hierarchical model's dispersion structure, calibration, and sensitivity to the generation-interval assumption were each checked explicitly; a shared (non-province-specific) dispersion parameter was retained on the basis of negligible predictive difference (PSIS-LOO), the model achieved 95.0% pooled 95% posterior-predictive interval coverage, and the province ranking was unchanged across a generation-interval sensitivity grid (Spearman {rho} = 1.0). Aggregation masks meaningful heterogeneity in transmission intensity at every spatial resolution examined; response prioritisation based on a single national or even provincial Rt risks directing attention away from the specific zones where transmission remains supercritical.
Li, J.; Yao, Q.; Pei, S.; Ning, N.
Show abstract
Antimicrobial-resistant organisms (AMROs) impose a major burden on healthcare systems, yet routine surveillance cannot readily distinguish colonization imported at admission from transmission acquired within hospitals. This gap is especially consequential because both processes may vary substantially across wards, while asymptomatic carriage, incomplete testing, imperfect diagnostic sensitivity, and patient movement obscure the underlying transmission dynamics. To address this challenge, we developed a blockwise agent-based iterated filter (BAIF) for inference in a patient-level transmission model on a dynamic ward co-location network. The model tracks susceptible and colonized patients as they move across wards, represents unobserved colonization histories, and incorporates the recorded testing schedule and imperfect diagnostic sensitivity. BAIF uses blockwise likelihood evaluation and resampling to estimate ward-block-specific transmission rates and importation probabilities in this high-dimensional latent system. Synthetic experiments showed that BAIF recovered these parameters from partially observed outbreaks. We then applied the framework to hospitalization and microbiological surveillance data collected from 2012 to 2016 at an urban quaternary care hospital in New York City for four AMROs. Transmission and importation were highly heterogeneous across ward blocks. Elevated transmission was repeatedly concentrated in the same ward groups, whereas blocks with the highest importation varied by pathogen. By distinguishing importation-dominated from transmission-dominated ward blocks, the framework can inform more targeted surveillance and infection-control strategies. More broadly, BAIF provides an effective inference framework for high-dimensional, partially observed agent-based models on dynamic contact networks.
Verheyden, J. G. L.; Mudogo, C. N.; Jacquet, W.
Show abstract
Background The 2026 Bundibugyo virus disease outbreak in the Democratic Republic of the Congo became the largest recorded epidemic caused by Bundibugyo virus and showed unusually rapid early growth. Initial projections warned that it could become one of the largest Ebola-family outbreaks on record, but these were based on assumed mortality totals and intervention scenarios rather than repeated comparison with subsequently observed surveillance data. We assessed how accurately the outbreak could have been forecast in real time from routinely published national situation reports, whether a simple ensemble improved on its component models, and how forecasts should inform operational capacity planning. Methods We conducted a retrospective pseudo-prospective rolling-origin study using daily cumulative confirmed cases and deaths reported from 14 May to 20 July 2026 (68 calendar days; 61 numeric reports). Each date with a newly reported national total became an eligible forecast origin once seven numeric observations were available, yielding 55 origins. At each origin, all later data were withheld. Two prespecified models were refitted using only information then available: a recent seven-increment baseline and a Bayesian negative-binomial surveillance-maturity model fitted by Markov Chain Monte Carlo. A Gompertz model was fitted identically as a benchmark. Forecasts were produced for 7, 14, and 21 days and evaluated only when an observation existed on the exact target date. Bayesian forecasts were prospectively recalibrated using only previously realised errors from forecast-maturity-qualified origins. We also evaluated a baseline-Bayesian ensemble, with horizon- and outcome-specific weights selected prospectively by an expanding-origin procedure using an asymmetric operational loss function. Findings A statistically and externally corroborated surveillance-maturity discontinuity occurred on 28 May 2026 (robust z score 13{middle dot}0). The Gompertz model had the poorest 80% predictive-interval coverage at every horizon and for both outcomes (25{middle dot}9-34{middle dot}1% for cases; 0{middle dot}0-23{middle dot}5% for deaths) and was excluded. Recalibration corrected consistent Bayesian underprediction and was applied at 24 of 55 origins for 21-day forecasts. It substantially improved case coverage (7-day, 60{middle dot}0% to 80{middle dot}0%; 14-day, 41{middle dot}0% to 76{middle dot}9%) and improved death forecasting in both accuracy and calibration (7-day median absolute percentage error, 10{middle dot}3% to 7{middle dot}7%; 80% coverage, 33{middle dot}3% to 95{middle dot}6%). Ensemble weights differed by outcome: case forecasts were baseline-heavy (w{approx}0{middle dot}8-0{middle dot}9), whereas death forecasts were near parity (w{approx}0{middle dot}5-0{middle dot}6). For deaths, the ensemble improved mean weighted interval score over both component models at every horizon (7-day: baseline 24{middle dot}8, Bayesian 20{middle dot}5, ensemble 17{middle dot}9). Interpretation Routine situation-report data supported useful short-term forecasting, but no single model was best on every criterion. Three forms of maturity shaped forecast reliability: epidemic, surveillance, and forecast maturity. Forecast maturity was the most robust and model-independent finding and supports 21 days, rather than 28 days, as the longest routine operational horizon. The study provides a reproducible and adjustable framework for combining simple and complex models and communicating uncertainty to operational decision-makers.
Verheyden, J. G. L.; Mudogo, C. N.
Show abstract
Anticipating which health zone will report the next confirmed case is operationally distinct from forecasting national case counts and matters for prepositioning response capacity; most spatial spread models rely on mobile-phone mobility data unavailable in the Democratic Republic of the Congo (DRC). We modelled the discrete-time hazard of a first reported confirmed case across 106 health zones in four provinces affected by the 2026 Bundibugyo virus disease outbreak (47 affected, 59 at risk, 26 July 2026), comparing four connectivity specifications,none, road-distance, a gravity score, and an incidence-weighted force-of-infection (FOI) term, fitted within an identical Bayesian hierarchical hazard architecture. Evaluation used a rolling-origin design, cluster bootstrap resampling, leave-one-origin-out and non-overlapping-origin checks, and a kernel-parameter sensitivity grid, with top-10 hit rate the pre-specified primary metric, matched to the operational question of which few zones warrant attention; AUC-PR, top-5 hit rate, and median rank percentile were secondary. FOI had the highest top-10 hit rate (42.6%), approaching conventional significance against road-distance and no-connectivity comparators. On AUC-PR, a model with no connectivity term performed as well as or better than any connectivity specification (0.437 vs. 0.409 for FOI), a discrepancy we report rather than omit. Rankings were stable across the sensitivity grid (Spearman; 0.90-0.99) and across robustness checks. An incidence-weighted connectivity term modestly and specifically improves identification of the highest-risk zones, concentrated in top-k ranking rather than uniform across metrics. The evaluation is pseudo-prospective, since historical data-vintage snapshots could not rule out retrospective revision, pending verification via a pre-registered top-20 ranking. Keywords: Bundibugyo virus disease; Ebola; spatial epidemiology; hazard model; Bayesian statistics; Democratic Republic of the Congo; disease surveillance
Janies, D.; Guirales-Medrano, S.; Santos Silva, A. C. D.; Ford, C. T.
Show abstract
Cyclospora cayetanensis causes significant seasonal foodborne illness, characterized by recent large-scale nationwide outbreaks. We are currently experiencing the largest recorded outbreak in USA history during the summer of 2026. The summer 2026 outbreak emphasizes the urgent need for enhanced surveillance of infectious diseases. In this study, we use phylogenetic network analysis via the StrainHub framework to reconstruct the historical spread of C. cayetanensis in the USA using only mitochondrial sequence data and place of isolation metadata collected from 1997 to 2022. By quantifying parasite importation to localities through the metric indegree centrality, we identify Texas as a primary sink for C. cayetanensis introduction. Our finding aligns with historical epidemiologic investigations such as the 2013 multistate outbreak. While our model demonstrates the efficacy of reconstructing transmission pathways from molecular data alone, it also reveals critical surveillance gaps, particularly the masking of geographic origins in public datasets and the current absence of molecular data for the 2026 event. Moving forward, integrating real-time molecular surveillance with robust network analysis is essential to shift from reactive to proactive intervention strategies. We conclude that improved data sharing and reporting across food, clinical, and farm sectors are vital for disrupting transmission cycles of parasites. This work will protect public health and ensure a fresh clean supply of key foods for the nutrition of Americans.
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
Lokonon, B. E.; Haydon, D. T.; Fakas, C.; Bonfoh, B.
Show abstract
Background. Increasing evidence indicates that Ebola virus disease (EVD) survivors can remain a source of infection long after clinical recovery. Confirmed survivor-associated transmission events and genomic evidence linking the 2021 Guinea outbreak to viral lineages from the 2013-2016 West African epidemic have demonstrated that persistent infection in survivors can contribute to post-epidemic re-emergence. However, the population-level conditions under which survivor reservoirs may sustain recrudescence remain poorly understood. Methods. We developed an age-structured Bayesian transmission model to quantify survivor-driven recrudescence risk using historical Ebola outbreak data (1976-2022) and empirical viral persistence data from male survivors. Age-specific viral clearance probabilities were estimated for three age groups (less or equal to 25, 26-35, and >35 years). The recrudescence reproduction number (Rc) was derived using the next-generation matrix approach. Sensitivity analyses examined alternative assumptions regarding viral clearance and the potential contribution of female survivors. Results. The posterior mean recrudescence reproduction number remained below the persistence threshold (Rc=1) across all viral-clearance scenarios under the assumption of no female survivor contribution. Only by assuming the slowest rate of viral clearance and maximal female survivor contribution did the posterior mean for Rc exceed one (1.052; 95% CrI: 0.428-2.229), suggesting that survivor-driven transmission alone is unlikely to sustain Ebola re-emergence given our current understanding of recrudescence dynamics. Simulations showed that survivor-driven outbreak pressure (rate of survivor-initiated outbreaks) was driven primarily by outbreak size and clustering. Outbreaks involving less or equal to 5,000 EVD cases generally produced outbreak pressure below the estimated natural spillover rate, whereas outbreaks comparable in size to the 2013-2016 West African epidemic generated transient survivor-driven outbreak rates up to 7.8-fold higher than the natural spillover rate before declining to comparable levels within 3-7 years. Moreover, across all viral-clearance scenarios, older (>35 years) male survivors consistently exhibited the longest effective persistence durations and made the largest contribution to the recrudescence reproduction number. Conclusions. The human survivor reservoir represents a plausible complementary pathway for Ebola re-emergence, particularly following large epidemics and should be considered alongside zoonotic spillover as an important source of future outbreaks. Age-dependent viral clearance strongly shapes recrudescence dynamics, with older survivors contributing disproportionately to transmission potential. These findings support age-stratified survivor monitoring, extended persistence surveillance, and improved characterization of viral persistence in both male and female survivors to strengthen post-epidemic preparedness.
Schulz, S.; Rincon Hidalgo, A.; Jarynowski, A. K.; Zambrano, M.; Suer, J.; Thampi, A.; Ferretti, L.; Phuong, H. T.; Xu, C.; Mikolajczyk, R.; Pastor, R.; Jaeger, V. K.; Karch, A.; Belik, V.
Show abstract
Mass gathering events (MGEs) play a critical role for infectious disease dynamics on a population level as they provide opportunities for superspreading; however, underlying mechanisms remain insufficiently understood. We analyzed nationwide GPS-based, individual-level location data from mobile phone users in Germany between April and August 2024 with 16m spatial precision. Potentially infectious contacts were inferred from close co-location and linked to contact settings using OpenStreetMap data. Various MGEs, including EURO 2024 matches, major concerts, festivals, and fairs were compared using a common contact metric. Non-football events generated substantially more contacts than football events. While overall national contact numbers remained stable, MGEs produced so-called "small-world" contacts which gather people from distant locations into close proximity and could strongly enhance infectious disease dynamics. Crucially, most high-risk contacts occurred within two hours before the event, not at the event itself, and concentrated in public transport, leisure, and event-adjacent areas. Our work provides the first systematic and comparative evaluation of contact exposure across various types of MGEs and contact settings. Event-type-specific dynamics, particularly indirect and mobility-driven contacts, critically shape infection risk. These insights can inform accurate transmission modeling, targeted intervention and event-management strategies.
Bracher, J.; Wolffram, D.; Amaral Lind, R.; Bardeck, N.; Boehm, M.; Contreras, S.; Doenges, P.; Guenther, F.; Kaiser, R.; van de Kassteele, J.; Kuhlmann, A.; Lange, B.; Nemcova, B.; Priesemann, V.; Reinacher, U.; Rodiah, I.; Sandmann, F.; the RESPINOW Study Group, ; Schienle, M.
Show abstract
Respiratory diseases cause considerable morbidity in autumn and winter and are a priority in public health monitoring. In Germany, they are subject to a number of surveillance systems, including both pathogen-specific and syndromic indicators. In this paper we present a collaborative multi-target and multi-model real-time forecasting system rolled out during the 2024/25 season, and discuss differences to earlier efforts carried out during the COVID-19 pandemic. A total of nine models were run to generate forecasts of general practitioner consultations for acute respiratory infections (ARI), hospitalizations for severe acute respiratory infections (SARI) and confirmed cases of seasonal influenza and RSV. As all indicators were subject to retrospective revisions, forecasting models were combined with a nowcasting step. Whenever multiple models were available for the same indicator, we combined them into an ensemble. Nowcasts showed convincing performance, even though for some models Christmas break effects led to an upward bias in early January. Forecasts were overall well-calibrated and most models outperformed simple benchmark models. These improvements were generally more substantial for age-stratified than pooled targets, and concentrated at lead times of two to three weeks. Anticipating the peak timing and magnitude proved to be challenging, with many models predicting too flat curves with a too early turnaround (e.g. already in late January rather than mid-February for SARI). The combined ensemble forecast was among the best-performing approaches, but unlike in previous related projects did not consistently outperform individual models. We conclude by discussing learnings on the organization of collaborative forecasting projects in post-COVID-19 times and the potential of AI-supported modelling.
Weidemueller, P. H.; Esquivel Gomez, L. R.; Rodriguez-Barraquer, I.; Mueller, N. F.
Show abstract
Tracking how an infectious disease spreads in time and space relies on several distinct sources of surveillance data, reported case counts, viral concentrations in wastewater, seroprevalence surveys, and pathogen genomic sequences, each of which is imperfect and captures only part of the underlying transmission process. These data streams are typically analyzed separately or with highly parameterized, disease-specific models, making it difficult to combine their complementary strengths. Here we present MASCOT-DataStreams (MASCOT-DS), a BEAST2 software package that extends the structured coalescent model MASCOT to jointly infer prevalence over time and transmission rates between locations from any combination of case counts, wastewater concentrations, seroprevalence surveys, and pathogen phylogenies. Using simulated outbreaks in structured populations, we show that MASCOT-DS accurately recovers true prevalence trajectories and between-location migration rates. We then apply MASCOT-DS to genomic, case count, wastewater, and seroprevalence data from the SARS-CoV-2 Epsilon wave (winter 2020-21) in three San Francisco Bay Area counties, reconstructing county-level prevalence dynamics and quantifying transmission within and into the region. By systematically removing individual data streams, we find that genomic data are uniquely required to estimate transmission between locations, while seroprevalence data are essential for anchoring the overall magnitude of an outbreak; case counts and wastewater concentrations play largely interchangeable roles in capturing outbreak shape. These results demonstrate that integrating complementary epidemiological data streams substantially increases the certainty of transmission dynamics estimates compared to relying on any single data stream, and provides a framework for evaluating the added value of different surveillance strategies.
de Blasio, B. F.; Scalia Tomba, G.
Show abstract
ABSTRACT Patient movements within and between hospitals create networks that can facilitate the spread of nosocomial infections. Many relevant pathogens have long-lasting carriage, allowing colonised patients to move through multiple wards across successive admissions, thereby indirectly linking wards. Yet system-wide ward-level analyses based on individual patient trajectories remain rare, even though they can reveal important features for understanding and simulating the system. We analysed ward- and hospital-level networks using individual patient trajectories from the Norwegian Patient Registry (3.6 million registrations), covering hospital care for ~55% of the population over one year (2012). We characterised the global network structure and, at ward level, calculated multiple centralisation measures and assessed percolation-based connectivity. From these, we identified central wards using directed K-core decomposition combined with Gaussian mixture modelling clustering. Analyses were performed separately for inpatients and all patients, and extended by linking episodes across increasing time gaps to explore how assumptions about episode continuity influence inferred connectivity. Various patient-based statistics were also analysed. Patient movements generated sparse, regionally structured networks with clear core-periphery organisation. Inflow K-cores were larger than outflow cores, with central wards dominated by major hospital referral specialities and medical wards acting as key entry and exit points. Community structure differed by patient type: all-patient networks were largely locally contained, whereas inpatient networks spanned hospitals and regions. Temporal linking increased hub dominance only in all-patient networks, while inpatient structure remained comparatively stable. Dynamic diffusion simulations and a real outbreak analysis independently supported the identified structural backbone, demonstrating that central network hubs also represent the dominant potential pathways of patient-mediated spread under simplified transmission assumptions. The results reveal a robust hierarchical organisational structure of patient movements that can help prioritise surveillance, while underscoring the need for pathogen-specific epidemiological and contextual data for predictive modelling.