Back

Establishing wastewater-based SARS-CoV-2 variant surveillance independent of clinical isolates

Kociurzynski, R.; Reuter, S.; Donker, T.

2026-08-12 epidemiology
10.64898/2026.08.11.26359873 medRxiv
Show abstract

The COVID-19 pandemic remains a global concern, partly due to the rapid mutation rate of SARS-CoV-2 and the emergence of new variants. Wastewater surveillance has proven effective in estimating infection incidence and detecting variants earlier than clinical testing. Its importance has grown as testing rates decline due to milder disease progression. However, current methods typically rely on the prior classification of SARS-CoV-2 lineages or their signature mutations, which may delay detection. We present an alternative method that identifies changes in the viral genetic population over time without requiring prior lineage classification. This population-based approach was applied to sequencing data from wastewater samples, which are generally noisier than clinical samples. We analyzed publicly available sequencing samples from wastewater plants covering Swiss catchments in Altenrhein, St. Gall, Geneva, and Zurich. To address noise, only samples with read depths above 40 and genome coverage of at least 90% were included. Genetic diversity within pooled populations over two time periods was compared to assess changes in viral composition. We demonstrate that SARS-CoV-2 variants can be detected in wastewater sequencing data without prior lineage classification. Our method successfully detected shifts in genetic populations that corresponded to the emergence of known variants of concern (VOCs) in the analyzed regions. Notably, it also revealed the rising prevalence during the first surges of the Omicron variant. Despite the increased noise in wastewater compared to clinical samples, our approach remains effective. However, achieving reliable predictions depends on high sequencing depth, broad genome coverage, and frequent sampling.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.