Tracking SARS-CoV-2 genomic variants in wastewater sequencing data with LolliPop
Dreifuss, D.; Topolsky, I.; Icer Baykal, P.; Beerenwinkel, N.
Show abstract
During the COVID-19 pandemic, wastewater-based epidemiology has progressively taken a central role as a pathogen surveillance tool. Tracking viral loads and variant outbreaks in sewage offers advantages over clinical surveillance methods by providing unbiased estimates and enabling early detection. However, wastewater-based epidemiology poses new computational research questions that need to be solved in order for this approach to be implemented broadly and successfully. Here, we address the variant deconvolution problem, where we aim to estimate the relative abundances of genomic variants from next-generation sequencing data of a mixed wastewater sample. We introduce LolliPop, a computational method to solve the variant deconvolution problem. LolliPop is tailored to wastewater time series sequencing data and applies temporal regularization in the form of a fused ridge penalty. We show that this regularization is equivalent to kernel smoothing and that it makes abundance estimates robust to very high levels of missing data, which is common for wastewater sequencing. We use the bootstrap to produce confidence intervals, and develop analytical standard errors that can produce similar confidence intervals at a fraction of the computational cost. We demonstrate the application of our method to data from the Swiss wastewater surveillance efforts as well as on simulated data. Author SummaryWastewater-based epidemiology has become a valuable tool for tracking viruses like SARS-CoV-2 across entire communities. Sequencing wastewater can reveal which viral variants are circulating, offering early and unbiased insights into variant dynamics. A central challenge is to infer the relative abundances of these variants from observed mutation data. This task is complicated by the fact that variant profiles can be highly similar, and the data is often noisy with many missing values, especially when the incidence of the pathogen is low. We developed LolliPop, a statistical method that leverages the time series structure of wastewater data to robustly deconvolve variant abundances and compute fast confidence intervals. Using both simulated data and real data from the Swiss national variant monitoring, we show that LolliPop is accurate and robust to high levels of missing data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Linear-regression-based algorithms can succeed at identifying microbial functional groups despite the nonlinearity of ecological function 95%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 95%
- SCRaPL: hierarchical Bayesian modelling of associations in single cell multi-omics data 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.