Correcting Modification-Mediated Errors in Nanopore Sequencing by Nucleotide Demodification and in silico Correction
Chiou, C.-S.; Chen, B.-H.; Wang, Y.-W.; Kuo, N.-T.; Chang, C.-H.; Huang, Y.-T.
Show abstract
The accuracy of Oxford Nanopore Technology (ONT) sequencing has significantly improved thanks to new flowcells, sequencing kits, and basecalling algorithms. However, novel modifications untrained in the basecalling models can seriously reduce the quality. This paper reports a set of ONT-sequenced genomes with unexpected low quality ([~]Q30) due to extensive new modifications. Demodification by whole-genome amplification (WGA) significantly improved the quality of all genomes ([~]Q50-60) while losing the epigenome. We developed a computational method, Modpolish, for correcting modification-mediated errors without WGA. Modpolish produced high-quality genomes and uncovered the underlying modification motifs without loss of epigenome. Our results suggested that novel modifications are prone to ONT errors, which are correctable by WGA or Modpolish without additional short-read sequencing.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Accurate bacterial outbreak tracing with Oxford Nanopore sequencing and reduction of methylation-induced errors 95%
- Genealogical inference and more flexible sequence clustering using iterative PopPUNK 94%
- Taxor: Fast and space-efficient taxonomic classification of long reads with hierarchicalinterleaved XOR filters 94%
Similar papers in this journal
- GeneMates: an R package for Detecting Horizontal Gene Co-transfer between Bacteria Using Gene-gene Associations Controlled for Population Structure 95%
- MicroPIPE: An end-to-end solution for high-quality complete bacterial genome construction 95%
- Mining ancient microbiomes using selective enrichment of damaged DNA molecules 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.