Predicting COVID-19 incidence from seroprevalence and population-based cohort data using interpretable machine learning with differential privacy analysis
Krepel, J.; Binkyte, R.; Kerkouche, R.; Harries, M.; Klett-Tammen, C. J.; Fritz, M.; Kesselheim, S.; Kuehn, M.; Bazarova, A.; Lange, B.
Show abstract
During the COVID-19 pandemic, reported incidence data played a central role in public health surveillance and in tracking epidemic dynamics, although they provide limited insight into the behavioral, immunological, and socioeconomic drivers of transmission.Population-based seroprevalence studies with linked survey data offer a rich but untapped source of individual-level information that can complement routine surveillance. In this study, we investigate whether aggregated seroprevalence cohort data can be leveraged to predict local COVID-19 incidence and to identify interpretable predictors associated with transmission dynamics. Using data from the Multilocal SeroPrevalence (MuSPAD) study in Germany (2020--2022), we trained multiple machine learning models, including least absolute shrinkage and selection operator (LASSO), vector autoregressive models (VAR), multilayer perceptrons (MLPs), and long short-term memory neural networks (LSTMs), to predict location-specific seven-day incidence rates. Feature importance was assessed using regression coefficients where applicable and model-agnostic explainability methods, including Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). Across model classes, cohort-derived features enabled accurate prediction of local incidence, with time-aware models achieving the strongest performance. Consistent predictors included prior infection and testing history, employment-related changes, vaccination status, and mask-wearing behavior, highlighting the importance of behavioral and reporting-related signals. While differential privacy introduced modest degradation in predictive performance under strict privacy budgets, SHAP-based explanations remained stable, and LIME-based explanations were more sensitive to privacy-induced noise. These results demonstrate that aggregated cohort data encode meaningful and interpretable signals of population-level transmission dynamics. Population-based serosurveys therefore provide a complementary source of information for predicting local COVID-19 incidence and identifying key drivers of transmission beyond routine surveillance data. Our findings show that integrating interpretable machine learning with privacy-aware analysis enables actionable insights from sensitive cohort data, supporting their use in digital epidemiology and informing data-driven public health decision-making.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A prospective real-time transfer learning approach to estimate Influenza hospitalizations with limited data 95%
- Fast and Trustworthy Nowcasting of Dengue Fever: A Case Study Using Attention-Based Probabilistic Neural Networks in Sao Paulo, Brazil 95%
- Assessing the utility of COVID-19 case reports as a leading indicator for hospitalization forecasting in the United States 94%
Similar papers in this journal
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 96%
- Generalising uncertainty improves accuracy and safety of deep learning analytics applied to oncology 95%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 94%
Similar papers in this journal
Similar papers in this journal
- An adaptive weight ensemble approach to forecast influenza activity in the context of irregular seasonality 95%
- Short-term forecasting of COVID-19 in Germany and Poland during the second wave – a preregistered study 94%
- Permutation-based Identification of Important Biomarkers for Complex Diseases via Black-box Models 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.