How low can you go? Short-read polishing of Oxford Nanopore bacterial genome assemblies
Bouras, G.; Judd, L. M.; Edwards, R. A.; Vreugde, S.; Stinear, T. P.; Wick, R. R.
Show abstract
It is now possible to assemble near-perfect bacterial genomes using Oxford Nanopore Technologies (ONT) long reads, but short-read polishing is still required for perfection. However, the effect of short-read depth on polishing performance is not well understood. Here, we introduce Pypolca (with default and careful parameters) and Polypolish v0.6.0 (with a new careful parameter). We then show that: (1) all polishers other than Pypolca-careful, Polypolish-default and Polypolish-careful commonly introduce false-positive errors at low depth; (2) most of the benefit of short-read polishing occurs by 25x depth; (3) Polypolish-careful never introduces false-positive errors at any depth; and (4) Pypolca-careful is the single most effective polisher. Overall, we recommend the following polishing strategies: Polypolish-careful alone when depth is very low (<5x), Polypolish-careful and Pypolca-careful when depth is low (5-25x), and Polypolish-default and Pypolca-careful when depth is sufficient (>25x). Data SummaryPypolca is open-source and freely available on Bioconda, PyPI, and GitHub (github.com/gbouras13/pypolca). Polypolish is open-source and freely available on Bioconda and GitHub (github.com/rrwick/Polypolish). All code and data required to reproduce analyses and figures are available at github.com/gbouras13/depth_vs_polishing_analysis. All FASTQ sequencing reads are available at BioProject PRJNA1042815. A detailed list of accessions can be found in Table S1.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.