Identification and Masking of Artefactual and Misleading Within-Host Variants in Deep-Sequencing SARS-CoV-2 Data
Anker, K. M.; Hall, M.; Evans Pena, R.; Kemp, S. A.; Clarke, J.; Zhao, L.; Bonsall, D.; Grayson, N.; Bashton, M.; The COVID-19 Genomics UK (COG-UK) Consortium, ; Walker, A. S.; Golubchik, T.; Lythgoe, K.
Show abstract
Deep sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artefacts. In this study, we show that recurrent artefactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency (MAF) thresholds. Using data from the UKs Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artefacts are predominantly sequencing centre-rather than protocol-specific. Each centre exhibits a modest, distinct set of recurrent artefactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artefactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Together, these findings highlight the importance of explicit, dataset-aware artefact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens such as SARS-CoV-2.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- periscope: Sub-Genomic RNA Identification in SARS-CoV-2 Genomic Sequencing Data 96%
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 94%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 94%
Similar papers in this journal
- Limited genomic reconstruction of SARS-CoV-2 transmission history within local epidemiological clusters 95%
- The origins and molecular evolution of SARS-CoV-2 lineage B.1.1.7 in the UK 94%
- High Resolution analysis of Transmission Dynamics of Sars-Cov-2 in Two Major Hospital Outbreaks in South Africa Leveraging Intrahost Diversity 94%
Similar papers in this journal
- Analytical validity of nanopore sequencing for rapid SARS-CoV-2 genome analysis 94%
- SARS-CoV-2 within-host population expansion, diversification and adaptation in zoo tigers, lions and hyenas 93%
- Shotgun Transcriptome and Isothermal Profiling of SARS-CoV-2 Infection Reveals Unique Host Responses, Viral Diversification, and Drug Interactions 93%
Similar papers in this journal
Similar papers in this journal
- Recovery of deleted deep sequencing data sheds more light on the early Wuhan SARS-CoV-2 epidemic 93%
- quick analysis of sedimentary ancient DNA using quicksand 93%
- Genomic analysis of European Drosophila populations reveals major longitudinal structure, continent-wide selection, and unknown DNA viruses 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.