Viral genetic variability in wastewater predicts changes in community infection levels
Hill, D. T.; Schulman, R.; Caldas, I. V.; Dunham, C.; Zhu, Y.; Lamson, D. M.; Rickerman, L.; St. George, K.; Ahmed-Braimah, Y.; Green, H.; Kmush, B.; Middleton, F.; Larsen, D. A.
Show abstract
Sequencing viruses found in community wastewater facilitates the study of diversity in circulating viruses at the population level. By analyzing 12,290 wastewater samples collected between January 2023 and April 2025 in New York State, USA from 196 sampling sites across 57 counties, we assessed the diversity of the SARS-CoV-2 genome and how it changed over time compared to changes in COVID-19 infections and hospitalizations. We calculated three measures of SARS-CoV-2 genome diversity across all samples: nucleotide diversity ({pi}), Shannon diversity (H), and viral variant count. We found that diversity increased with a rise in COVID-19 incidence and hospitalizations for all three measures (with a Spearman{rho} > 0.8, p<0.001). The genetic diversity of the spike protein region had the highest correlation with the incidence of cases ({rho} = 0.92, p<0.001 for{pi} , {rho} = 0.91, p <0.001 for H), and the statewide count of virus variants had a correlation coefficient of{rho} = 0.85 (p<0.001) with case incidence. Additionally, the genetic diversity of the spike protein predicted 90.1 percent of the variance of COVID-19 case incidence. Our results demonstrate the potential for viral diversity analysis from wastewater in predicting epidemiological outcomes.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systematic SARS-CoV-2 S Gene Sequencing in Wastewater Samples Enables Early Lineage Detection and Uncovers Rare Mutations in Portugal 97%
- Development of a COVID-19 warning system for neighborhood-scale wastewater-based epidemiology in low incidence situations 97%
- Wastewater to clinical case (WC) ratio of COVID-19 identifies insufficient clinical testing, onset of new variants of concern and population immunity in urban communities 97%
Similar papers in this journal
- Metrics to relate COVID-19 wastewater data to clinical testing dynamics 98%
- Longitudinal metatranscriptomic sequencing of Southern California wastewater representing 16 million people from August 2020-21 reveals widespread transcription of antibiotic resistance genes. 97%
- Longitudinal SARS-CoV-2 RNA Wastewater Monitoring Across a Range of Scales Correlates with Total and Regional COVID-19 Burden in a Well-Defined Urban Population 97%
Similar papers in this journal
- Rapid, large-scale wastewater surveillance and automated reporting system enabled early detection of nearly 85% of COVID-19 cases on a University campus 96%
- High frequency, high throughput quantification of SARS-CoV-2 RNA in wastewater settled solids at eight publicly owned treatment works in Northern California shows strong association with COVID-19 incidence 96%
- Assessing multiplex tiling PCR sequencing approaches for detecting genomic variants of SARS-CoV-2 in municipal wastewater 95%
Similar papers in this journal
- Predictive power of wastewater for nowcasting infectious disease transmission: a retrospective case study of five sewershed areas in Louisville, Kentucky 94%
- How early into the outbreak can surveillance of SARS-CoV-2 in wastewater tell us? 94%
- Antibiotic resistance genes, antibiotic residues, and microplastics in influent and effluent wastewater from treatment plants in Norway, Iceland, and Finland 94%
Similar papers in this journal
- Nationwide trends in COVID-19 cases and SARS-CoV-2 wastewater concentrations in the United States 97%
- Evaluation of sampling frequency and normalization of SARS-CoV-2 wastewater concentrations for capturing COVID-19 burdens in the community 96%
- A High-Throughput Microfluidic Quantitative PCR Platform for the Simultaneous Quantification of Pathogens, Fecal Indicator Bacteria, and Microbial Source Tracking Markers 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.