Back

Harnessing the power of technical and natural variation in 116 yeast datasets to benchmark long read assembly pipelines

Gettle, N.; Gallone, B.; Verstrepen, K.; Stelkens, R.

2022-03-19 bioinformatics
10.1101/2022.03.17.484703 bioRxiv
Show abstract

With increases in throughput and reductions in cost, long read sequencing has become the standard for most genome assembly projects and has opened up new avenues for large-scale genomic research. While more amenable to assembly than short-read sequence data, long-read datasets tend to have higher error rates. To address this problem numerous tools have been developed to correct reads before assembly and polish assembled contigs. Although, numerous studies have been conducted to assess or benchmark these tools, few capture the real variance in long read sequence data that might affect tool performance much less full pipeline performance. To address these shortcomings, we compiled a dataset containing long-read sequences of 116 different strains of brewers yeast, S. cerevisiae, gathered largely from public databases and evaluated different assembly-related tools as well as their interactions. We found that pre-assembly short-read error correction of long reads combined with post-assembly short-read polishing provided the best assemblies. We also found that correction/polishing steps with uncorrected long reads often lead to degradation of assembly quality. Finally, we show which tools and pipelines work best with different types of input data.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.