Harnessing the power of technical and natural variation in 116 yeast datasets to benchmark long read assembly pipelines
Gettle, N.; Gallone, B.; Verstrepen, K.; Stelkens, R.
Show abstract
With increases in throughput and reductions in cost, long read sequencing has become the standard for most genome assembly projects and has opened up new avenues for large-scale genomic research. While more amenable to assembly than short-read sequence data, long-read datasets tend to have higher error rates. To address this problem numerous tools have been developed to correct reads before assembly and polish assembled contigs. Although, numerous studies have been conducted to assess or benchmark these tools, few capture the real variance in long read sequence data that might affect tool performance much less full pipeline performance. To address these shortcomings, we compiled a dataset containing long-read sequences of 116 different strains of brewers yeast, S. cerevisiae, gathered largely from public databases and evaluated different assembly-related tools as well as their interactions. We found that pre-assembly short-read error correction of long reads combined with post-assembly short-read polishing provided the best assemblies. We also found that correction/polishing steps with uncorrected long reads often lead to degradation of assembly quality. Finally, we show which tools and pipelines work best with different types of input data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Accuracy of de novo assembly of DNA sequences from double-digest libraries varies substantially among software 97%
- Chromosome-level hybrid de novo genome assemblies as an attainable option for non-model organisms 96%
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.