Fast and accurate assembly of Nanopore reads via progressive error correction and adaptive read selection
Chen, Y.; Nie, F.; Xie, S.-Q.; Zheng, Y.-F.; Bray, T.; Dai, Q.; Wang, Y.-X.; Huang, Z.-J.; Wang, D.-P.; He, L.-J.; Luo, F.; Wang, J.-X.; Liu, Y.-Z.; Xiao, C.-L.
Show abstract
Although long Nanopore reads are advantageous in de novo genome assembly, applying Nanopore reads in genomic studies is still hindered by their complex errors. Here, we developed NECAT, an error correction and de novo assembly tool designed to overcome complex errors in Nanopore reads. We proposed an adaptive read selection and two-step progressive method to quickly correct Nanopore reads to high accuracy. We introduced a two-stage assembler to utilize the full length of Nanopore reads. NECAT achieves superior performance in both error correction and de novo assembly of Nanopore reads. NECAT requires only 7,225 CPU hours to assemble a 35X coverage human genome and achieves a 2.28-fold improvement in NG50. Furthermore, our assembly of the human WERI cell line showed an NG50 of 29 Mbp. The high-quality assembly of Nanopore reads can significantly reduce false positives in structure variation detection.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 96%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
- Long-read and chromosome-scale assembly of the hexaploid wheat genome achieves high resolution for research and breeding 94%
Similar papers in this journal
Similar papers in this journal
- HISAT-3N: a rapid and accurate three-nucleotide sequence aligner 95%
- Ultra-low input single tube linked-read library method enables short-read NGS systems to generate highly accurate and economical long-range sequencing information for de novo genome assembly and haplotype phasing 95%
- Illumina But With Nanopore: Sequencing Illumina libraries at high accuracy on the ONT MinION using R2C2 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.