A Gold Standard Dataset for Lineage Abundance Estimation from Wastewater
Ferdous, J.; Kunkleman, S.; Taylor, W.; Harris, A.; Gibas, C. J.; Schlueter, J. A.
Show abstract
During the SARS-CoV-2 pandemic, genome-based wastewater surveillance sequencing has been a powerful tool for public health to monitor circulating and emerging viral variants. As a medium, wastewater is very complex because of its mixed matrix nature, which makes the deconvolution of wastewater samples more difficult. Here we introduce a gold standard dataset constructed from synthetic viral control mixtures of known composition, spiked into a wastewater RNA matrix and sequenced on the Oxford Nanopore Technologies platform. We compare the performance of eight of the most commonly used deconvolution tools in identifying SARS-CoV-2 variants present in these mixtures. The software evaluated was primarily chosen for its relevance to the CDC wastewater surveillance reporting protocol, which until recently employed a pipeline that incorporates results from four deconvolution methods: Freyja, kallisto, Kraken2/Bracken, and LCS. We also tested Lollipop, a deconvolution method used by the Swiss SARS-CoV2 Sequencing Consortium, and three recently-published methods: lineagespot, Alcov, and VaQuERo. We found that the commonly used software Freyja outperformed the other CDC pipeline tools in correct identification of lineages present in the control mixtures, and that the newer method VaQuERo was similarly accurate, with minor differences in the ability of the two methods to avoid false negatives and suppress false positives. These results provide insight into the effect of the tiling primer scheme and wastewater RNA extract matrix on viral sequencing and data deconvolution outcomes. HighlightsO_LIGeneration of a gold standard dataset C_LIO_LIComparative evaluation of relative abundance estimation software C_LIO_LIEvaluation of deconvolution methods used in CFSANs CWAP pipeline C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Implementing Building-Level SARS-CoV-2 Wastewater Surveillance on a University Campus 97%
- Applicability of Neighborhood and Building Scale Wastewater-Based Genomic Epidemiology to Track the SARS-CoV-2 Pandemic and other Pathogens 97%
- An alternative approach for bioanalytical assay development for wastewater-based epidemiology of SARS-CoV-2 97%
Similar papers in this journal
- Biomarkers Selection for Population Normalization in SARS-CoV-2 Wastewater-based Epidemiology 96%
- When Case Reporting Becomes Untenable: Can Sewer Networks Tell Us Where COVID-19 Transmission Occurs? 96%
- Highly efficient and sensitive membrane-based concentration process allows quantification, surveillance, and sequencing of viruses in large volumes of wastewater. 96%
Similar papers in this journal
- centriflaken: an automated data analysis pipeline for assembly and in silico analyses of foodborne pathogens from metagenomic samples 97%
- Wastewater surveillance in smaller college communities may aid future public health initiatives 96%
- Precision long-read metagenomics sequencing for food safety by detection and assembly of Shiga toxin-producing Escherichia coli in irrigation water 96%
Similar papers in this journal
- Machine-learning based detection of adventitious microbes in T-cell therapy cultures using long read sequencing 95%
- Wastewater genomic surveillance captures early detection of Omicron in Utah 95%
- Assessing Microbial Diversity in Soil Samples Along the Potomac River: Implications for Environmental Health 94%
Similar papers in this journal
- Quantitative Trend Analysis of SARS-CoV-2 RNA in Municipal Wastewater Exemplified with Sewershed-Specific COVID-19 Clinical Case Counts 97%
- Evaluation of sampling frequency and normalization of SARS-CoV-2 wastewater concentrations for capturing COVID-19 burdens in the community 96%
- A High-Throughput Microfluidic Quantitative PCR Platform for the Simultaneous Quantification of Pathogens, Fecal Indicator Bacteria, and Microbial Source Tracking Markers 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.