Deletion detection in SARS-CoV-2 genomes using multiplex-PCR sequencing from COVID-19 patients: elimination of false positives
Jiang, N.; Dewey, C. N.; Yin, J.
Show abstract
Deletions are prevalent in the genomes of SARS-CoV-2 isolates from COVID-19 patients, but their roles in the severity, transmission, and persistence of disease are poorly understood. Millions of COVID-19 swab samples from patients have been sequenced and made available online, offering an unprecedented opportunity to study such deletions. Multiplex PCR-based amplicon sequencing (amplicon-seq) has been the most widely used method for sequencing clinical COVID-19 samples. However, existing bioinformatics methods applied to negative control samples sequenced by multiplex-PCR sequencing often yield large numbers of false-positive deletions. We found that these false positives commonly occur in short alignments, at low frequency and depth, and near primer-binding sites used for whole-genome amplification. To address this issue, we developed a filtering strategy, validated with positive control samples containing a known deletion. Our strategy accurately detected the known deletion and removed more than 99% of false positives. This method, applied to public COVID-19 swab data, revealed that deletions occurring independently of transcription regulatory sequences were about 20-fold less common than previously reported; however, they remain more frequent in symptomatic patients. Our optimized approach should enhance the reliability of SARS-CoV-2 deletion characterization from surveillance studies. Finally, our approach may guide the development of more reliable bioinformatics pipelines for genome sequence analyses of other viruses.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-Throughput, Single-Copy Sequencing Reveals SARS-CoV-2 Spike Variants Coincident with Mounting Humoral Immunity during Acute COVID-19 96%
- High-resolution analysis of Merkel Cell Polyomavirus in Merkel Cell Carcinoma reveals distinct integration patterns and suggests NHEJ and MMBIR as underlying mechanisms 94%
- Detecting SARS-CoV-2 Cryptic Lineages using Publicly Available Whole Genome Wastewater Sequencing Data 94%
Similar papers in this journal
- Alternate primers for whole-genome SARS-CoV-2 sequencing 95%
- Long-read transcriptomics of Ostreid herpesvirus 1 uncovers a conserved expression strategy for the capsid maturation module and pinpoints a mechanism for evasion of the ADAR-based antiviral defence 95%
- Optimized SMRT-UMI protocol produces highly accurate sequence datasets from diverse populations - application to HIV-1 quasispecies 94%
Similar papers in this journal
- A Short Plus Long-Amplicon Based Sequencing Approach Improves Genomic Coverage and Variant Detection In the SARS-CoV-2 Genome 97%
- Nanopore Sequencing of SARS-CoV-2: Comparison of Short and Long PCR-tiling Amplicon Protocols 96%
- Oligonucleotide Capture Sequencing of the SARS-CoV-2 Genome and Subgenomic Fragments from COVID-19 Individuals 95%
Similar papers in this journal
- Targeted Hybridization Capture of SARS-CoV-2 and Metagenomics Enables Genetic Variant Discovery and Nasal Microbiome Insights 96%
- Leptospira transcriptome sequencing using long-read technology reveals unannotated transcripts and potential polyadenylation of mRNA molecules 93%
- Integrase-associated niche differentiation of endogenous large DNA viruses in crustaceans 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.