SARS-CoV-2 sequencing artifacts associated with targeted PCR enrichment and read mapping
Ellegaard, K. M.; Gunalan, V.; Sieber, R.; Baig, S. J.; Larsen, N. B.; Bennedbaek, M.; Bybjerg-Grauholm, J.; Escobar-Herrera, L. A.; Hansen, T. N. G.; Thorsen, T. H.; Krusager, A.; Aasbjerg, G. N.; Al-Tamimi, N. S.; Westergaard, C.; Svarrer, C. W.; Rasmussen, M.; Stegger, M.
Show abstract
Protocols and pipelines for SARS-CoV-2 genome sequencing were rapidly established when the COVID-19 outbreak was declared a pandemic. The most widely used approach for sequencing SARS-CoV-2 includes targeted enrichment by PCR, followed by shotgun sequencing and reference-based genome assembly. As the continued surveillance of SARS-CoV-2 worldwide is transitioning towards a lower level of intensity, it is timely to re-visit the sequencing protocols and pipelines established during the acute phase of the pandemic. In the current study, we have investigated the impact of primer scheme and reference genome choice by sequencing samples with multiple primer schemes (Artic V3, V4.1 and V5.3.2) and re-processing reads with multiple reference genomes. We have also analysed the temporal development in ambiguous base calls during the emergence of the BA.2.86.x variant. We found that the primers used for targeted enrichment can result in recurrent ambiguous base calls, which can accumulate rapidly in response to the emergence of a new variant. We also found examples of consistent base calling errors, associated with PCR artifacts and amplicon drop-out. Similarly, misalignments and partially mapped reads on the reference genome resulted in ambiguous base calls, as well as defining mutations being omitted from the assembly. These findings highlight some key limitations of using targeted enrichment by PCR and reference-based genome assembly for sequencing SARS-CoV-2, and the importance of continuously monitoring and updating primer schemes and bioinformatic pipelines.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis of the ARTIC V4 and V4.1 SARS-CoV-2 primers and their impact on the detection of Omicron BA.1 and BA.2 lineage defining mutations 96%
- Bioinformatic investigation of discordant sequence data for SARS-CoV-2: insights for robust genomic analysis during pandemic surveillance 96%
- Comparison of R9.4.1/Kit10 and R10/Kit12 Oxford Nanopore flowcells and chemistries in bacterial genome reconstruction 95%
Similar papers in this journal
- Nanopore Sequencing of SARS-CoV-2: Comparison of Short and Long PCR-tiling Amplicon Protocols 97%
- Performance of amplicon and capture based next-generation sequencing approaches for the epidemiological surveillance of Omicron SARS-CoV-2 and other variants of concern. 96%
- Oligonucleotide Capture Sequencing of the SARS-CoV-2 Genome and Subgenomic Fragments from COVID-19 Individuals 95%
Similar papers in this journal
- A Rapid, Cost-Effective Tailed Amplicon Method for Sequencing SARS-CoV-2 95%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
- MicroPIPE: An end-to-end solution for high-quality complete bacterial genome construction 94%
Similar papers in this journal
- Rapid, High-Throughput, Cost Effective Whole Genome Sequencing of SARS-CoV-2 Using a Condensed One Hour Library Preparation of the Illumina DNA Prep Kit 96%
- SARS-CoV-2 Genome Sequencing Methods Differ In Their Ability To Detect Variants From Low Viral Load Samples 96%
- Accurate and Reproducible Whole-Genome Genotyping for Bacterial Genomic Surveillance with Nanopore Sequencing Data 95%
Similar papers in this journal
- Rapid detection of G6PD deficiency SNPs using a novel amplicon-based MinION Sequencing Assay 94%
- Optimised multiplex amplicon sequencing for mutation identification using the MinION nanopore sequencer 94%
- VarLOCK - sequencing independent, rapid detection of SARS-CoV-2 variants of concern for point-of-care testing, qPCR pipelines and national wastewater surveillance 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.