Genomic sequence characteristics and the empiric accuracy of short-read sequencing
Marin, M. G.; Vargas, R.; Harris, M.; Jeffrey, B.; Epperson, L. E.; Durbin, D.; Strong, M.; Salfinger, M.; Iqbal, Z.; Akhundova, I.; Vashakidze, S.; Crudu, V.; Rosenthal, A.; Farhat, M. R.
Show abstract
BackgroundShort-read whole genome sequencing (WGS) is a vital tool for clinical applications and basic research. Genetic divergence from the reference genome, repetitive sequences, and sequencing bias, reduce the performance of variant calling using short-read alignment, but the loss in recall and specificity has not been adequately characterized. For the clonal pathogen Mycobacterium tuberculosis (Mtb), researchers frequently exclude 10.7% of the genome believed to be repetitive and prone to erroneous variant calls. To benchmark short-read variant calling, we used 36 diverse clinical Mtb isolates dually sequenced with Illumina short-reads and PacBio long-reads. We systematically study the short-read variant calling accuracy and the influence of sequence uniqueness, reference bias, and GC content. [a] ResultsReference based Illumina variant calling had a recall [≥]89.0% and precision [≥]98.5% across parameters evaluated. The best balance between precision and recall was achieved by tuning the mapping quality (MQ) threshold, i.e. confidence of the read mapping (recall 85.8%, precision 99.1% at MQ [≥] 40). Masking repetitive sequence content is an alternative conservative approach to variant calling that maintains high precision (recall 70.2%, precision 99.6% at MQ[≥]40). Of the genomic positions typically excluded for Mtb, 68% are accurately called using Illumina WGS including 52 of the 168 PE/PPE genes (34.5%). We present a refined list of low confidence regions and examine the largest sources of variant calling error. ConclusionsOur improved approach to variant calling has broad implications for the use of WGS in the study of Mtb biology, inference of transmission in public health surveillance systems, and more generally for WGS applications in other organisms.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- SNPPar: identifying convergent evolution and other homoplasies from microbial whole-genome alignments 95%
- Phylogenomic and genomic analysis reveals unique and shared genetic signatures of Mycobacterium kansasii complex species 95%
- Tracking SARS-CoV-2 variants of concern in wastewater: an assessment of nine computational tools using simulated genomic data 95%
Similar papers in this journal
- Identifying the essential genes of Mycobacterium avium subsp. hominissuis with Tn-Seq using a rank-based filter procedure. 94%
- A program for real-time surveillance of SARS-CoV-2 genetics 93%
- A needle in a haystack: metagenomic DNA sequencing to quantify Mycobacterium tuberculosis DNA and diagnose tuberculosis 93%
Similar papers in this journal
- High precision Neisseria gonorrhoeae variant and antimicrobial resistance calling from metagenomic Nanopore sequencing 96%
- Ultra-low input single tube linked-read library method enables short-read NGS systems to generate highly accurate and economical long-range sequencing information for de novo genome assembly and haplotype phasing 94%
- periscope: Sub-Genomic RNA Identification in SARS-CoV-2 Genomic Sequencing Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.