CovSyn: an agent-based model for synthesizing COVID-19 course of disease and contact tracing data
Wu, Y.-H.; Nordling, T. E. M.
Show abstract
The COVID-19 pandemic has demonstrated the shortcomings of epidemiological modelling for guiding policy decisions. Moreover, the modelling efforts resulted in many models yielding different predictions, creating a need to compare these predictions to determine which model is most accurate. We introduce a data synthesis algorithm, CovSyn, designed to generate synthetic COVID-19 datasets providing sufficiently detailed information for benchmarking epidemiological models against a known synthetic ground truth. CovSyn utilizes observed infections with contact tracing, testing, course of disease data, and a contact network based on municipality statistics, which categorises connections into household, school, workplace, healthcare, and municipality. The models initial parameters and boundaries are derived from empirical data, including the first community outbreak of COVID-19 in Taiwan and clinical observations. Comprehensive parameter space exploration for optimal results is done by the Firefly algorithm. We demonstrate it and validate our estimates by comparing state transition times, daily social contacts, and associated secondary attack rates against a structured dataset and clinical observations. Our simulations align with prior research and this dataset. Most state transition times from 10,000 simulations are within uncertainty ranges. Daily contact numbers and their distribution across layers match empirical findings. Our model accurately reproduced the first COVID-19 outbreak in Taiwan, achieving high accuracy with observed cumulative confirmed cases (R2 = 0.9) across daily, 7-day moving average, and 31-day moving average levels. Each synthetic subject contains demographic data (age, gender, occupation), course of disease (latent/incubation periods, testing, isolation, critical illness, recovery, and death dates), and contact network data including daily interactions with infected and uninfected individuals. Our algorithm offers a valid alternative for developing and benchmarking epidemiological models to advance COVID-19 forecasting research. Author summaryWe present a novel approach to enhance the testing of epidemiological infection modelling and demonstrate spread prediction accuracy by testing COVID-19 simulation models on Taiwanese synthetic data with known ground truth. CovSyn generates comprehensive synthetic data at the individual level, encompassing demographic characteristics (age, gender, occupation), course of disease (infection dates, symptom onset, recovery dates), and contact tracing information (daily social interactions including household, school, workplace, healthcare, and municipality). It can provide a consistent, standardized synthetic dataset for model evaluation, addressing the previous challenge of comparing COVID-19 models that used disparate data sources and different time periods. In this study, we detail the algorithm and demonstrate its reliability by creating a synthetic dataset for the first outbreak of SARS-CoV-2 in Taiwan and comparing it with the collected Taiwan COVID-19 dataset. Future work will focus on benchmarking state-of-the-art forecasting models using our synthetic data. Through CovSyns detailed individual-level data, we aim to advance the development of more accurate epidemiological models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An integrated framework for building trustworthy data-driven epidemiological models: Application to the COVID-19 outbreak in New York City 98%
- BharatSim: An agent-based modelling framework for India 97%
- Estimation of heterogeneous instantaneous reproduction numbers with application to characterize SARS-CoV-2 transmission in Massachusetts counties 96%
Similar papers in this journal
Similar papers in this journal
- Assessing the effects of non-pharmaceutical interventions on SARS-CoV-2 transmission in Belgium by means of an extended SEIQRD model and public mobility data 96%
- Modeling the early phase of the Belgian COVID-19 epidemic using a stochastic compartmental model and studying its implied future trajectories 95%
- Using an Agent-Based Model to Assess K-12 School Reopenings Under Different COVID-19 Spread Scenarios – United States, School Year 2020/21 95%
Similar papers in this journal
- Extended compartmental model for modeling COVID-19 epidemic in Slovenia 97%
- Modeling the Effect of Lockdown Timing as a COVID-19 Control Measure in Countries with Differing Social Contacts 97%
- Tracing contacts to evaluate the transmission of COVID-19 from highly exposed individuals in public transportation 96%
Similar papers in this journal
- Deep reinforcement learning framework for controlling infectious disease outbreaks in the context of multi-jurisdictions 96%
- A Machine Learning Approach to Differentiate Between COVID-19 and Influenza Infection Using Synthetic Infection and Immune Response Data 96%
- Individual-based modeling of COVID-19 transmission in college communities 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.