Back

Epidemics

Elsevier BV

All preprints, ranked by how well they match Epidemics's content profile, based on 116 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Understanding RSV Resurgence Following COVID-19 in Ontario, Canada: Evaluating the Roles of Contact Patterns and Maternal Immunity

Parpia, A.; Wright, J.; Gharouni, A.; Thampi, N.; Fitzpatrick, T.

2026-08-31 epidemiology 10.64898/2026.08.28.26361657 medRxiv
Top 0.1%
52.7%
Show abstract

Background: Respiratory syncytial virus (RSV) remains a leading cause of hospitalization in infancy, with severe outcomes influenced by both contact patterns and passive immunity. Non-pharmaceutical interventions (NPIs) during the COVID-19 pandemic suppressed RSV circulation and reduced opportunities for maternal immune boosting, potentially altering protection among newborns. We evaluated whether incorporating time-varying maternal immunity improves the ability of an age-structured transmission model to predict post-pandemic RSV hospitalization patterns in infants. Methods: We analyzed population-based RSV hospitalizations among Ontario (Canada) infants (<1 year) from July 2, 2017 to June 25, 2024, using linked administrative databases. A deterministic compartmental model across seven age classes was calibrated against pre-pandemic data using Latin Hypercube Sampling. We compared a model incorporating time-varying contact rates alone against a specification that additionally included time-varying maternal immunity. Results: Both specifications accurately reproduced pre-pandemic seasonality and macro-level post-pandemic resurgence features. The constant maternal immunity model showed slightly better accuracy in capturing the 2021/22 peak compared to the time-varying maternal immunity specification. However, both qualitatively captured the continued near-absence of RSV and the observed peak was captured within the 95% credible intervals. While both models precisely captured the timing and overwhelming surge of admissions that occurred in 2022/23, they failed to capture the premature peak timing and magnitude in 2023/24. Conclusions: Incorporating time-varying maternal immunity did not improve model accuracy post-pandemic. While maternal protection is essential for evaluating infant immunizations, population-level contact shifts primarily shaped post-pandemic RSV seasonality, indicating that models must account for these mechanisms of RSV transmission dynamics.

2
Short- and long-term causes of West Nile virus risk in Europe: a spatiotemporal model accounting for under-reporting

Bastard, J.; Assaad, C.; Marti, R.; Tran, A.; Metras, R.; DURAND, B.

2026-08-23 infectious diseases 10.64898/2026.08.20.26360902 medRxiv
Top 0.1%
39.7%
Show abstract

Models that provide risk maps for zoonoses often lack (i) a spatiotemporal autocorrelation component, yet crucial in understanding the spread of infectious diseases, (ii) accounting for heterogeneity in case reporting, and (iii) a causal framework for explanatory variables. Here, we addressed these limitations with a model system, West Nile virus, a vector-borne pathogen transmitted in a bird reservoir, and affecting humans and horses. We built a spatiotemporal occupancy model and fitted it to notified (human and horse) case data. Based on a directed acyclic graph, we estimated the causal effects of conjectural weather variables (i.e. changing in the short-term) vs. structural variables (i.e. changing in the long-term) on WNV circulation in the bird reservoir, besides assessing variables associated with case reporting. By computing population attributable fractions, we found the contribution of conjectural weather variables to WNV outbreaks in Europe to be globally higher than the structure of the bird community.

3
Quantifying infection-relevant contact patterns among young children in childcare settings in the United States.

Loo, S. L.; Nande, A.; Hill, A. L.; Truelove, S.

2026-08-21 epidemiology 10.64898/2026.08.18.26360623 medRxiv
Top 0.1%
39.7%
Show abstract

Age is a primary determinant of symptom severity and transmission patterns for many infectious diseases, motivating the use of age-stratified models parameterized by contact matrices. In the United States, the absence of direct contact surveys has required estimating synthetic contact matrices from demographic data on household size, school attendance, and workforce participation. However, this likely underestimates contacts among children under age 5, who often attend group childcare missing from censuses. The goal of this study was to use nationally-representative data on childcare arrangements (the Early Childhood Program Participation Survey) to reconstruct daily contacts occurring in childcare settings, and augment existing all-age contact matrices. For infants under 1 year of age, we estimated 0.2 daily contacts with other infants, increasing to 0.7 daily contacts with same-age peers for 1- or 2-year-olds, 1.3 for 3-year-olds, and 3.5 for 4-year-olds. Including childcare settings increases estimated contacts among young children by up to six fold. Using simulations of measles outbreaks in inadequately vaccinated populations, we show that prior contact matrices significantly underestimated the outbreak frequency, size, and impact on preschool age groups. Our findings highlight the need for targeted data collection on childcare contacts to improve model-based evaluation of interventions particularly for young children.

4
Sub-epidemic model forecasts for COVID-19 pandemic spread in the USA and European hotspots, February-May 2020

Chowell, G.; Rothenberg, R.; Roosa, K.; Tariq, A.; Hyman, J. M.; Luo, R.

2020-07-04 epidemiology 10.1101/2020.07.03.20146159 medRxiv
Top 0.1%
38.8%
Show abstract

Mathematical models have been widely used to understand the dynamics of the ongoing coronavirus disease 2019 (COVID-19) pandemic as well as to predict future trends and assess intervention strategies. The asynchronicity of infection patterns during this pandemic illustrates the need for models that can capture dynamics beyond a single-peak trajectory to forecast the worldwide spread and for the spread within nations and within other sub-regions at various geographic scales. Here, we demonstrate a five-parameter sub-epidemic wave modeling framework that provides a simple characterization of unfolding trajectories of COVID-19 epidemics that are progressing across the world at different spatial scales. We calibrate the model to daily reported COVID-19 incidence data to generate six sequential weekly forecasts for five European countries and five hotspot states within the United States. The sub-epidemic approach captures the rise to an initial peak followed by a wide range of post-peak behavior, ranging from a typical decline to a steady incidence level to repeated small waves for sub-epidemic outbreaks. We show that the sub-epidemic model outperforms a three-parameter Richards model, in terms of calibration and forecasting performance, and yields excellent short- and intermediate-term forecasts that are not attainable with other single-peak transmission models of similar complexity. Overall, this approach predicts that a relaxation of social distancing measures would result in continuing sub-epidemics and ongoing endemic transmission. We illustrate how this view of the epidemic could help data scientists and policymakers better understand and predict the underlying transmission dynamics of COVID-19, as early detection of potential sub-epidemics can inform model-based decisions for tighter distancing controls.

5
Failure to balance social contact matrices can bias models of infectious disease transmission

Hamilton, M. A.; Knight, J.; Mishra, S.

2022-07-31 infectious diseases 10.1101/2022.07.28.22278155 medRxiv
Top 0.1%
38.6%
Show abstract

Spread of transmissible diseases is dependent on contact patterns in a population (i.e. who contacts whom). Therefore, many epidemic models incorporate contact patterns within a population through contact matrices. Social contact survey data are commonly used to generate contact matrices; however, the resulting matrices are often imbalanced, such that the total number of contacts reported by group A with group B do not match those reported by group B with group A. While the importance of balancing contact matrices has been acknowledged, how these imbalances affect modelled projections (e.g., peak infection incidence, impact of public health measures) has yet to be quantified. Here, we explored how imbalanced contact matrices from age-stratified populations (<15, 15+) may bias transmission dynamics of infectious diseases. First, we compared the basic reproduction number of an infectious disease when using imbalanced versus balanced contact matrices from 177 demographic settings. Then, we constructed a susceptible exposed infected recovered transmission model of SARS-CoV-2 and compared the influence of imbalanced matrices on infection dynamics in three demographic settings. Finally, we compared the impact of age-specific vaccination strategies when modelled with imbalanced versus balanced matrices. Models with imbalanced matrices consistently underestimated the basic reproduction number, had delayed timing of peak infection incidence, and underestimated the magnitude of peak infection incidence. Imbalanced matrices also influenced cumulative infections observed per age group, and the projected impact of age-specific vaccination strategies. For example, when vaccine was prioritized to individuals <15 in a context where individuals 15+ underestimated their contacts with <15, imbalanced models underestimated cumulative infections averted among 15+ by 24.4%. We conclude stratified transmission models that do not consider reciprocity of contacts can generate biased projections of epidemic trajectory and impact of targeted public health interventions. Therefore, modellers should ensure and report on balancing of their contact matrices for stratified transmission models. AUTHOR SUMMARYTransmissible diseases such as COVID-19 spread according to who contacts whom. Therefore, mathematical transmission models - used to project epidemics of infectious diseases and assess the impact of public health interventions - require estimates of who contacts whom (also referred to as a contact matrix). Contact matrices are commonly generated using contact surveys, but this data is often imbalanced, where the total number of contacts reported by group A with group B does not match those reported by group B with group A. Although these imbalances have been acknowledged as an issue, the influence of imbalanced matrices on modelled projections (e.g. peak incidence, impact of public health interventions) has not been explored. Using a theoretical model of COVID-19 with two age groups (<15 and 15+), we show models with imbalanced matrices had biased epidemic projections. Models with imbalanced matrices underestimated the initial spread of COVID-19 (i.e. the basic reproduction number), had later time to peak COVID-19 incidence and smaller peak COVID-19 incidence. Imbalanced matrices also influenced cumulative infections observed per age group, and the estimated impact of an age-specific vaccination strategy. Given imbalanced contact matrices can reshape transmission dynamics and model projections, modellers should ensure and report on balancing of contact matrices.

6
A bootstrap particle filter for viral Rt inference and forecasting using wastewater data

Xiao, W. F.; Wang, Y.; Goel, N.; Wolfe, M.; Koelle, K.

2026-03-06 epidemiology 10.64898/2026.03.06.26347747 medRxiv
Top 0.1%
35.2%
Show abstract

Wastewater is increasingly being recognized as an important data stream that can contribute to infectious disease surveillance and forecasting. With this recognition, a growing number of statistical inference approaches are being developed to use wastewater data to provide quantitative insights into epidemiological dynamics. However, few existing approaches have allowed for systematic integration of data streams for inference, for example by combining case incidence data and/or serological data with wastewater data. Furthermore, only a subset of existing approaches have been able to handle missing data without imputation and to handle datasets with different sampling times or intervals. Here, we develop a statistically rigorous, yet lightweight, approach to infer and forecast time-varying effective reproduction numbers (Rt values) using longitudinal wastewater virus concentrations either alone or jointly with additional data streams including case incidence data and serological data. Our approach relies on a state-space modeling approach for inference and forecasting, within the context of a simple bootstrap particle filter. We first describe the structure of our underlying disease transmission process model as well as our observation models. Using a mock dataset, we then show that Rt can be accurately estimated by interfacing this model with case incidence data, wastewater data, or a combination of these two data streams using the bootstrap particle filter. Of note, we show that these data streams alone do not allow for reconstruction of underlying infection dynamics due to structural parameter unidentifiability. We then apply our particle filter to a previously analyzed SARS-CoV-2 dataset from Zurich that includes case data and wastewater data. Our analyses of these real-world datasets indicate that incorporation of process noise (in the form of environmental stochasticity) into the state space model greatly improves our ability to reconstruct the latent variables of the model. We further show that underlying infection dynamics can be made identifiable through the incorporation of serological data and that the bootstrap particle filter can be used to make forecasts of Rt, case incidence, and wastewater virus concentrations. We hope that the inference approach presented here will lead to greater reliance on wastewater data for disease surveillance and forecasting that will aid public health practitioners in responding to infectious disease threats.

7
Empirical contact networks reveal heterogeneous outbreak risks in a UK long-term care facility: a modelling study

Pi, L.; Davis, E. L.; Danon, L.; Hollingsworth, D.

2026-07-10 infectious diseases 10.64898/2026.07.07.26357238 medRxiv
Top 0.1%
34.4%
Show abstract

Long-term care facilities (LTCs) worldwide experienced disproportionately high infection and mortality rates during the COVID-19 pandemic, where essential care limits opportunities for contact segregation. However, empirical contact data remain scarce, limiting our understanding of how individual contact behaviours shape transmission in these settings. In this study, we developed a stochastic network-based transmission model parameterised using real-world self-reported contact data collected from a median-sized UK LTC unit. By incorporating high-resolution observational data that reflect routine care delivery patterns, we quantified how heterogeneity in contact networks influences outbreak dynamics. We found substantial variation in contact behaviour between individuals, resulting in highly heterogeneous transmission outcomes. Outbreak occurrence, timing, final size, and the likelihood of super-spreading events all varied markedly depending on the structure of the underlying contact network and the characteristics of the index case. Individuals with high contact activity were considerably more likely to initiate large outbreaks than those with fewer contacts. For a per-contact transmission probability of 10%, introduction of infection through the most highly connected individuals resulted in a greater than 75% probability of a large outbreak. Our findings indicate that preventing infection introduction through both residents and staff is critical for outbreak control in LTCs. Individuals with high contact activity were consistently associated with a greater probability of initiating large outbreaks, highlighting the importance of accounting for contact heterogeneity when designing surveillance and infection-control measures. More broadly, this study demonstrates the importance of accounting for contact-network heterogeneity when designing infection prevention and control measures in LTC settings, and highlights the value of integrating empirical contact data with transmission modelling to inform evidence-based outbreak preparedness, targeted surveillance, and infection-control strategies in long-term care facilities.

8
Population-Level Associations in the Spread of Co-Circulating Respiratory Viruses: A Multi-Method Statistical Investigation Using Incidence Data

Barth, N.; Carstens, G.; Kozanli, E.; Han, W.; Hermans, L.; Paolotti, D.; Abrams, S.; Molenberghs, G.; Hens, N.; Faes, C.; Eggink, D.; van Hoek, A. J.; Torneri, A.

2025-11-22 infectious diseases 10.1101/2025.11.20.25340550 medRxiv
Top 0.1%
34.3%
Show abstract

Respiratory infections remain a major global health burden, causing substantial morbidity and mortality worldwide. The responsible viruses circulate concurrently, potentially affecting each others dynamics, yet the extent and direction of such interactions remain poorly understood. Characterising these cross-pathogen effects at the population level is essential for elucidating transmission dynamics and guiding mitigation strategies. Using incidence data from a participatory syndromic surveillance system with multiplex PCR confirmation of specific pathogens, we applied complementary statistical approaches, including multivariate regression, endemic-epidemic, and distributed-lag models, to characterise immediate and delayed associations among seven major respiratory diseases. We show that these pathogens form a connected system in which some, such as SARS-CoV-2 and human seasonal coronaviruses, enhance each others transmission, whereas others, notably influenza, inhibit the concurrent circulation of competitors such as rhinovirus or parainfluenza virus. Effects were often directional rather than reciprocal: for instance, rhinovirus inhibited human seasonal coronaviruses but not vice versa, while mutual enhancement between human metapneumovirus and parainfluenza virus appeared across several models. Interaction patterns were time-dependent yet largely consistent, indicating persistent ecological interference among co-circulating respiratory viruses. By integrating multiple analytic frameworks, our study provides a comprehensive, data-driven view of how respiratory viruses coexist and compete, offering crucial insights for improved epidemic forecasting and mitigation strategies.

9
Temporal contact patterns and the implications for predicting superspreaders and planning of targeted outbreak control

Pung, R.; Firth, J. A.; Russell, T.; Rogers, T.; Lee, V. J.; Kucharski, A. J.

2023-12-11 epidemiology 10.1101/2023.11.22.23298919 medRxiv
Top 0.1%
34.0%
Show abstract

Epidemic models often heavily simplify the dynamics of human-to-human contacts, but the resulting bias in outbreak dynamics - and hence requirements for control measures - remains unclear. Even if high-resolution temporal contact data were routinely used for modelling, the role of this temporal network structure towards outbreak control is not well characterised. We address this by assessing dynamic networks across varied social settings in three ways. Firstly, we characterised the distribution of retained contacts over consecutive timesteps by developing a novel metric, the "retention index", which accounts for the change in the number of contacts over consecutive timesteps on a normalised scale with the extremes representing fully static and fully dynamic networks. Secondly, we described the repetition of contacts over the days by estimating the frequency of contact pairs occurring over the study duration. Thirdly, we distinguish the difference between superspreader and infectious individuals driving superspreading events by estimating the connectivity of an individual (i.e. individual has high connectivity in a timestep if he accounts for 80% of the contacts in the timestep) and the frequency of exhibiting high connectivity. Using 11 networks from 5 settings studied over 3-10 days, we estimated that more than 80% of the individuals in most settings were highly connected for only short periods. This suggests a challenge to identify superspreaders, and more individuals would need to be targeted as part of outbreak interventions to achieve the same reduction in transmission as predicted from a static network. Taking into account repeated contacts over multiple days, we estimated simple resource planning models might overestimate the number of contacts made by an infector by 20%-70%. In workplaces and schools, contacts in the same department accounted for most of the retained contacts. Hence, outbreak control measures would be better off targeting specific sub-populations in these settings to reduce transmission. In contrast, no obvious type of contact dominated the retained contacts in hospitals, so reducing the risk of disease introduction is critical to avoid disrupting the interdependent work functions. This study identified key epidemiological properties of temporal networks that potentially shape outbreak dynamics and illustrated the need for incorporating such properties in outbreak simulations. SignificanceDirectly transmitted infectious diseases spread through social contacts that can change over time. Modelling studies have largely focused on simplifying these contact patterns to predict outbreaks but the assumptions on contact patterns may bias results and, in turn, conclusions on the effectiveness of control measures. An ongoing challenge is, therefore, how to measure key properties of complex and dynamic networks to facilitate the development of network disease simulation models, which ensures that outbreak analysis is transparent and interpretable in the real-world context. To address this challenge, we analysed 11 networks from 5 different settings and developed new metrics to capture crucial epidemiological features of these networks. We showed that there is an inherent difficulty in identifying individual superspreaders reliably in most networks. In addition, the key types of individuals driving transmission vary across settings, thus requiring different outbreak control measures to reduce disease transmission or the risk of introduction. Simple models to mimic disease transmission in temporal networks may not capture the repeated contacts over the days, and hence could incorrectly estimate the resources required for outbreak control. Our study characterised temporal network data in epidemiologically relevant ways and is a step towards developing simplified contact networks to capture real-world contact patterns for future outbreak simulation studies.

10
Evaluating the use of social contact data to produce age-specific forecasts of SARS-CoV-2 incidence

Munday, J. D.; Abbott, S.; Meakin, S.; Funk, S.

2022-12-03 epidemiology 10.1101/2022.12.02.22282935 medRxiv
Top 0.1%
33.9%
Show abstract

Short-term forecasts can provide predictions of how an epidemic will change in the near future and form a central part of outbreak mitigation and control. Renewal-equation based models are increasingly popular. They infer key epidemiological parameters from historical epidemiological data and forecast future epidemic dynamics without requiring complex mechanistic assumptions. However, these models typically ignore interaction between age-groups, partly due to challenges in parameterising a time varying interaction matrix. Social contact data collected regularly by the CoMix survey during the COVID-19 epidemic in England, provide a means to inform interaction between age-groups in real-time. We developed an age-specific forecasting framework and applied it to two age-stratified time-series: incidence of SARS-CoV-2 infection, estimated from a national infection and antibody prevalence survey; and, reported cases according to the UK national COVID-19 dashboard. Jointly fitting our model to social contact data from the CoMix study, we inferred a time-varying next generation matrix which we used to project infections and cases in the four weeks following each of 29 forecast dates between October 2021 and November 2022. We evaluated the forecasts using proper scoring rules and compared performance with three other models with alternative data and specifications alongside two naive baseline models. Overall, incorporating age-interaction improved forecasts of infections and the CoMix-data-informed model was the best performing model at time horizons between two and four weeks. However, this was not true when forecasting cases. We found that age-group-interaction was most important for predicting cases in children and older adults. The contact-data-informed models performed best during the winter months of 2020 - 2021, but performed comparatively poorly in other periods. We highlight challenges regarding the incorporation of contact data in forecasting and offer proposals as to how to extend and adapt our approach, which may lead to more successful forecasts in future.

11
Modelling patterns of SARS-CoV-2 circulation in the Netherlands, August 2020-February 2022, revealed by a nationwide sewage surveillance program

van Boven, M.; Hetebrij, W. A.; Swart, A. M.; Nagelkerke, E.; van der Beek, R. F.; Stouten, S.; Hoogeveen, R. T.; Miura, F.; Kloosterman, A.; van der Drift, A.-M. R.; Welling, A.; Lodder, W. J.; de Roda Husman, A. M.

2022-05-30 infectious diseases 10.1101/2022.05.25.22275569 medRxiv
Top 0.1%
32.7%
Show abstract

BackgroundSurveillance of SARS-CoV-2 in wastewater offers an unbiased and near real-time tool to track circulation of SARS-CoV-2 at a local scale, next to other epidemic indicators such as hospital admissions and test data. However, individual measurements of SARS-CoV-2 in sewage are noisy, inherently variable, and can be left-censored. AimWe aimed to infer latent virus loads in a comprehensive sewage surveillance program that includes all sewage treatment plants (STPs) in the Netherlands and covers 99.6% of the Dutch population. MethodsA multilevel Bayesian penalized spline model was developed and applied to estimate time- and STP-specific virus loads based on water flow adjusted SARS-CoV-2 qRT-PCR data from 1-4 sewage samples per week for each of the >300 STPs. ResultsThe model provided an adequate fit to the data and captured the epidemic upsurges and downturns in the Netherlands, despite substantial day-to-day measurement variation. Estimated STP virus loads varied by more than two orders of magnitude, from approximately 1012 (virus particles per 100,000 persons per day) in the epidemic trough in August 2020 to almost 1015 in many STPs in January 2022. Epidemics at the local levels were slightly shifted between STPs and municipalities, which resulted in less pronounced peaks and troughs at the national level. ConclusionAlthough substantial day-to-day variation is observed in virus load measurements, wastewater-based surveillance of SARS-CoV-2 can track long-term epidemic progression at a local scale in near real-time, especially at high sampling frequency.

12
Quantifying superspreading in bacterial STI outbreaks using phylodynamics

Sevilla, J.; Kende, J.; Duchene, S.; Meehan, M. T.

2026-08-17 epidemiology 10.64898/2026.08.14.26360404 medRxiv
Top 0.1%
31.9%
Show abstract

Bacterial sexually transmitted infections (STIs) pose a major global public health challenge, with Neisseria gonorrhoeae being of particular concern due to its persistently high prevalence and increasing antimicrobial resistance. The emergence of multidrug-resistant strains has narrowed treatment options, highlighting the importance of prevention. In this context, knowing whether there is superspreading (transmission heterogeneity) within a population becomes crucial for accurate public health measures. However, classic methods to quantify superspreading rely on dense contact tracing, and this is not always feasible. As an alternative, we can use Bayesian phylodynamic modelling to infer transmission dynamics, including superspreading. Yet modelling transmission dynamics using bacterial data remains problematic, although it is widely used for viral data. Here, we apply a multi-type birth-death model parametrised to quantify superspreading in N. gonorrhoeae outbreaks, estimating the fraction and relative impact of superspreaders and reproductive numbers for superspreaders and non-superspreaders. We also use a hierarchical modelling strategy with partial pooling to increase the power for detecting superspreading in each cluster. Model performance was successfully evaluated across a range of superspreading scenarios using both transmission-informed phylogenies and sequence data with phylogenetic uncertainty. Application to empirical genomic data revealed a substantial role of superspreading in N. gonorrhoeae transmission during the COVID-19 pandemic in Australia. These results highlight the impact of superspreading in N. gonorrhoeae transmission and the importance of detecting it to efficiently stop the dissemination of the disease

13
A comparison of random mixing in a structured agent-based model with empirical contact survey data

Suer, J.; Ponge, J.; Brüggemann, M.; Burgard, J. P.; Belik, V.; Hellingrath, B.; Hidalgo, A. R.; Jarynowski, A. K.; Pastor, R.; Phuong, H. T.; Schulz, S.; Thampi, A.; Xu, C.; Zambrano, M.; Mikolajczyk, R.; Karch, A.; Jaeger, V. K.

2025-09-22 epidemiology 10.1101/2025.09.18.25336044 medRxiv
Top 0.1%
31.8%
Show abstract

Agent-based models (ABMs) are powerful tools for simulating disease spread, relying on individual-level representations and interaction rules from which emergent dynamics arise. These rules need to be accurately specified as minor differences can lead to vastly different disease dynamics. An important component in ABMs is the contact behaviour. To decrease the computational complexity the contact behaviour is often assumed as random mixing within settings. Here, individuals randomly contact other individuals associated with the same settings, such as colleagues in their workplace, where the setting associations are based on empirical data such as census data. However, the validity of the random mixing assumption within settings remains unclear. We address this gap by comparing the contact structure in a large-scale ABM (GEMS) with empirical contact survey data (COVIMOD). We compare the age contact matrices for households, schools, workplaces, all remaining contact settings combined, and all contacts combined. This includes calculating the difference matrix and the sum of squared errors (SSE), i.e. the element-wise squared difference. Our results demonstrate that random mixing in settings generated based on known age-compositions like households (SSE:0.7(95%CI0.4-0.9)), schools (SSE:0.7(95%CI:0.3-1.1)) and workplaces (SSE:0.5(95%CI:0.2-0.7)), can capture the basic interaction patterns. However, it fails to account for age-related variations in contact numbers, leading to discrepancies between the simulated and observed contact behaviour. The largest differences arise for contacts outside of households, schools and workplaces (SSE:3.8(95%CI:1.2-6.5)), due to the models structure. These contacts are modelled as random regional contacts not capturing the age-structured behaviour observed in COVIMOD. We conclude that random mixing in accurately defined settings provides an approximation for contact structures in settings where the age-structure of associated individuals is similar to the observed contact structure. For settings where the age-structure deviates from the contact structure, advanced methods are required to represent real-world contact structures. Author SummaryInfectious disease modelling is a prominent tool for understanding disease spread and evaluating potential countermeasures. Agent-based models simulate the behaviour of individuals based on a set of simple rules and allow for a detailed representation of disease transmissions. One essential rule is the contact behaviour of individuals, which determines the potential transmission pathways of pathogens. Agent-based models often assume simple random-mixing of individuals in the locations they are associated with, such as their households or workplaces. We investigate if this assumption is appropriate by comparing the contact structure simulated in the agent-based model GEMS with the contact data gathered through the Germany-wide COVIMOD survey. We find that the random mixing can serve as an approximation for settings such as household, schools and workplaces where the contact structure is similar to the age structure of the individuals present in the setting. However, discrepancies arise as random mixing cannot account for the differing number of contacts by age or household size. We identified the most discrepancies for contacts outside of the household, school and workplace. Here, the highly age-structured contact behaviour observed in COVIMOD strongly deviates from the random mixing of all ages observed in GEMS.

14
Artificial Intelligence And Computational Methods For Modelling And Forecasting Influenza And Influenza-Like Illness: A Scoping Review

Onifade, I.; Adeoye, A.; Bayode, M.; Michael, I.; Akangbe, B.; Akomolafe, O.; Ajisafe, T.; Hossain, D.; Owoeye, O.

2025-05-06 epidemiology 10.1101/2025.05.05.25327004 medRxiv
Top 0.1%
31.6%
Show abstract

BackgroundThe persistnt resurgence of influence and influenza-like illness despite concerted vaccination interventions is a global health burden, thus necessitating accurate tools for early intervention and preparedness. This scoping review aims to map the currently available literature on artificial intelligence (AI)-based forecasting models for seasonal influenza and to identify trends in those published models, approaches, and research gaps. MethodsA detailed search was conducted in PubMed, Scopus, and IEEE Xplore to find relevant studies published between 2014 and 2025. The AI techniques (such as machine learning and deep learning) applied in predicting seasonal influenza activity are considered eligible studies. Model types, data inputs, performance metrics, and validation approaches were summarized on data that were extracted and charted. ResultsNine studies met the inclusion criteria and were included. Owing to their effectiveness in solving temporal sequence models, many deep learning models have been applied, including the long short-term memory (LSTM) model and the CNN LSTM hybrid model. The data sources are epidemiological records, meteorological variables and social media signals. Most of the models achieved excellent predictive accuracy, but shortcomings in model interpretability, external validation or consistency across performance reporting became issues. ConclusionsAlthough AI-based models show promising capabilities for predicting influenza, there are still issues related to standardization and deployment in the real world. Future work should focus on real-time data integration, external validation and interpretable transferable models appropriate for a wide variety of health settings. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=120 SRC="FIGDIR/small/25327004v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@10faed3org.highwire.dtl.DTLVardef@9ec211org.highwire.dtl.DTLVardef@d6ee82org.highwire.dtl.DTLVardef@c4d1a7_HPS_FORMAT_FIGEXP M_FIG C_FIG This graphical framework encapsulates AI-based forecasting models for seasonal influenza, depicted as a navigational chart through the research terrain. A central magnifying glass over a globe anchors the global health challenge, guiding the viewer through a flowchart-like journey. A funnel filters literature from PubMed, Scopus, and IEEE Xplore (2014-2025), yielding 9 pivotal studies. Layered icons delineate machine learning and deep learning models, with LSTM and CNN-LSTM hybrids highlighted. Interconnected circles symbolize diverse data inputs-- epidemiological, meteorological, and social media--converging into a data integration hub. The bar chart connotes high predictive accuracy, tempered by a warning sign flagging interpretability, validation, and reporting challenges. A roadmap at the journeys end points to future horizons: real-time data integration, external validation, and interpretable models, charting the course for advancing global influenza preparedness.

15
Co-circulating pathogens of humans: A systematic review of mechanistic transmission models

Shaw, K.; Peterson, J.; Jalali, N.; Ratnavale, S.; Alkuzweny, M.; Barbera, C.; Costello, A.; Emerick, L.; Espana, G.; Meyer, A.; Mowry, S.; Poterek, M.; de Souza Moreira, C.; Morgan, E.; Moore, S. M.; Perkins, A.

2024-09-16 infectious diseases 10.1101/2024.09.16.24313749 medRxiv
Top 0.1%
31.5%
Show abstract

Historically, most mathematical models of infectious disease dynamics have focused on a single pathogen, despite the ubiquity of co-circulating pathogens in the real world. We conducted a systematic review of 311 published papers that included a mechanistic, population-level model of co-circulating human pathogens. We identified the types of pathogens represented in this literature, techniques used, and motivations for conducting these studies. We also created a complexity index to quantify the degree to which co-circulating pathogen models diverged from single-pathogen models. We found that the emergence of new pathogens, such as HIV and SARS-CoV-2, precipitated modeling activity of the emerging pathogen with established pathogens. Pathogen characteristics also tended to drive modeling activity; for example, HIV suppresses the immune response, eliciting interesting dynamics when it is modeled with other pathogens. The motivations driving these studies were varied but could be divided into two major categories: exploration of dynamics and evaluation of interventions. Finally, we found that model complexity quickly increases as additional pathogens are added. Future potential avenues of research we identified include investigating the effects of misdiagnosis of clinically similar co-circulating pathogens and characterizing the impacts of one pathogen on public health resources available to curtail the spread of other pathogens.

16
Infectious disease modeling for public health practice: projections, scenarios, and uncertainty in three phases of outbreak response

Brouwer, A. F.; Eisenberg, M. C.; Dean, N. E.; Hochheiser, H.; Huang, P.; Coyle, J. R.; Rennert, L.

Top 0.1%
31.5%
Show abstract

Public health departments need evidence-backed scenario projections to support decision making in infectious disease outbreaks. However, traditional infectious disease models are often not readily deployable or responsive to the urgent questions and priorities of public health departments or health systems. Moreover, uncertainty in model outputs is not always adequately assessed or communicated, potentially undermining trust among public health practitioners and the public. To address these issues, we, the Insight Net Modeling Guidance for Public Health Working Group, used early COVID-19 data from Michigan to illustrate modeling approaches that can be used to answer urgent questions in three key phases of outbreak response: prior to local introduction, early exponential growth, and established transmission with potential interventions. In each phase, we integrate case, hospitalization, and death data and capture ranges of plausible future trajectories. These models, which produce status quo and scenario projections, are intended to inform planning and motivate action rather than forecast precise future outcomes. Importantly, this work offers guidance to focus modeling efforts and provides examples and code for how to fit and implement these models, ultimately serving as both a conceptual guide and practical toolkit to support more transparent, timely, and appropriate use of models in outbreak response.

17
Identifying leptospirosis hotspots in Fiji using a One Health model that incorporates watershed-scale pathogen transport

Wakwella, A. R.; Klein, C. J.; Woodberry, O.; Lau, C. L.; Wenger, A.; Jupiter, S. D.; Jenkins, A. P.; Mayfield, H. J.

2026-08-04 epidemiology 10.64898/2026.08.03.26359543 medRxiv
Top 0.1%
31.4%
Show abstract

Leptospirosis is a water-related zoonotic disease with complex transmission pathways, including direct transmission from infected animals and indirect transmission through contaminated soil and water. Identifying key areas to implement targeted infection prevention and control strategies is challenging, as a range of risk factors across different scales can drive human infection. We aimed to develop an epidemiological modelling approach to predict key transmission pathways driving leptospirosis infection, including risks ranging from household level factors to the movement of pathogens across watersheds. We combined a causal Bayesian network with a novel hydrological pathogen transport model to predict leptospirosis across Fiji and found key infection hotspots adjacent to rivers and within degraded watersheds; a dynamic overlooked by previous epidemiological models. We used a wide range of data for model parameterisation (e.g., expert elicitation, epidemiological surveys, and environmental data) and found that predictive validity improved when expert input on risk factors was used to guide model parameterisation, improving R2 for predicted versus observed seroprevalence from 0.71 to 0.91. Our One Health modelling approach can be used to support the design and evaluation of environment-based disease prevention strategies at national scales.

18
Characterizing spatiotemporal variation in transmission heterogeneity during the 2022 mpox outbreak in the USA

Love, J.; LaPrete, C. R.; Sheets, T. R.; Vega Yon, G. G.; Thomas, A.; Samore, M. H.; Keegan, L. T.; Adler, F. R.; Slayton, R. B.; Spicknall, I. H.; Toth, D. J.

2023-05-14 epidemiology 10.1101/2023.05.10.23289580 medRxiv
Top 0.1%
31.2%
Show abstract

Understanding how transmission heterogeneity varies over the course of an enduring infectious disease outbreak improves understanding of observed disease dynamics and informs public health strategy. We quantified the spatiotemporal variation in transmission heterogeneity for the 2022 mpox outbreak in the US using the dispersion parameter of the offspring distribution, k. Our methods fit negative binomial distributions to transmission chain offspring distributions informed by a large mpox contact tracing dataset. We found that estimates of transmission heterogeneity varied across the outbreak, but overall estimated transmission heterogeneity was low. When testing our methods on simulated data, estimate accuracy depended on contact tracing data accuracy and completeness. Because the actual contact tracing data had high incompleteness, the values of k estimated from the empirical data may therefore be artificially high. Through simulation, we explore a method to correct estimated k for data incompleteness and, further, explore baseline expectations for temporal dynamics of k.

19
From surveillance maturity to analytical readiness: an estimand-first framework for real-time outbreak analysis under imperfect data

Verheyden, J. G. L.; Mudogo, C. N.; Jacquet, W.

2026-08-14 epidemiology 10.64898/2026.08.12.26360299 medRxiv
Top 0.1%
31.2%
Show abstract

Background: Real-time outbreak analyses are often requested before surveillance systems have stabilised or epidemics have generated enough information for the desired inference. Existing approaches address surveillance quality, forecasting, estimands and identifiability separately, but do not provide a common rule for deciding which analytical product is supportable at a particular data vintage. We developed an estimand-first framework for analytical readiness. Methods: The framework distinguishes surveillance maturity (S), epidemic-process informativeness (E) and estimand-specific analytical readiness, defined as whether the available data vintage, observation process, method and decision-matched validation support a specified inference for a specified decision. We stress-tested four implications using longitudinal data from the 2018-2020 Ebola response in eastern Democratic Republic of the Congo (DRC), archived geographic forecasts, independent forecasting data from Western Area, Sierra Leone, and a targeted mortality-identifiability experiment. Results: During a documented DRC surveillance disruption and recovery, seven-day persistence forecasts had all-health-zone absolute errors of 1, 3, 18 and 4 cases across pre-shock, acute-shock, early-recovery and recovery origins; the largest error occurred during early recovery. Four-week reported-case trend multipliers changed from 0.67 and 0.73 to 1.12 and 1.29, while the final fit was strongly overdispersed (Pearson dispersion 7.65), demonstrating asynchronous readiness across estimands. Archived geographic forecasts improved a Top-3 allocation decision over cumulative burden at only one origin despite consistently lower Brier scores for one specification. In Western Area, persistence forecast mean absolute error increased from 39.1 cases at one week to 82.3 at four weeks, and a history-to-horizon ratio did not define a universal threshold. An observed reported case-fatality ratio of 0.40 was compatible with constructed latent fatality values from 0.10 to 0.80; increasing the reported denominator narrowed sampling uncertainty without reducing structural uncertainty. Conclusions: Analytical readiness is task- and vintage-specific rather than a property of a dataset. More data, model convergence or narrow intervals cannot substitute for estimand definition, observation-process awareness, decision-matched validation and explicit identification analysis. Keywords: outbreak analytics; surveillance maturity; analytical readiness; estimand; identifiability; forecasting; Ebola; reporting process; decision-matched validation

20
Predicting the impact of non-pharmaceutical interventions against COVID-19 on Mycoplasma pneumoniae in the United States

Park, S. W.; Noble, B.; Howerton, E.; Nielsen, B. F.; Jiudice, S. S.; Ambroggio, L.; Dominguez, S.; Messacar, K.; Grenfell, B.

2024-08-20 infectious diseases 10.1101/2024.08.19.24312254 medRxiv
Top 0.1%
31.2%
Show abstract

The introduction of non-pharmaceutical interventions (NPIs) against COVID-19 disrupted circulation of many respiratory pathogens and eventually caused large, delayed outbreaks, owing to the build up of the susceptible pool during the intervention period. In contrast to other common respiratory pathogens that re-emerged soon after the NPIs were lifted, longer delays (> 3 years) in the outbreaks of Mycoplasma pneumoniae (Mp), a bacterium commonly responsible for respiratory infections and pneumonia, have been reported in Europe and Asia. As Mp cases are continuing to increase in the US, predicting the size of an imminent outbreak is timely for public health agencies and decision makers. Here, we use simple mathematical models to provide robust predictions about a large upcoming Mp outbreak in the US. Our model further illustrates that NPIs and waning immunity are important factors in driving long delays in epidemic resurgence.