An efficient error correction and accurate assembly tool for noisy long reads
Hu, J.; Wang, Z.; Sun, Z.; Hu, B.; Ayoola, A. O.; Liang, F.; Li, J.; Sandoval, J. R.; Cooper, D. N.; Ye, K.; Ruan, J.; Xiao, C.-L.; Wang, D.; Wu, D.-D.; Wang, S.
Show abstract
Long read sequencing data, particularly those derived from the Oxford Nanopore (ONT) sequencing platform, tend to exhibit a high error rate. Here, we present NextDenovo, a highly efficient error correction and assembly tool for noisy long reads, which achieves a high level of accuracy in genome assembly. NextDenovo can rapidly correct reads; these corrected reads contain fewer errors than other comparable tools and are characterized by fewer chimeric alignments. We applied NextDenovo to the assembly of high quality reference genomes of 35 diverse humans from across the world using ONT Nanopore long read sequencing data. Based on these de novo genome assemblies, we were able to identify the landscape of segmental duplications and gene copy number variation in the modern human population. The use of the NextDenovo program should pave the way for population-scale long-read assembly, thereby facilitating the construction of human pan-genomes, using Nanopore long read sequencing data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 96%
- Long-read and chromosome-scale assembly of the hexaploid wheat genome achieves high resolution for research and breeding 95%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
Similar papers in this journal
- Highly accurate assembly polishing with DeepPolisher 97%
- Ultra-low input single tube linked-read library method enables short-read NGS systems to generate highly accurate and economical long-range sequencing information for de novo genome assembly and haplotype phasing 96%
- Decoil: Reconstructing extrachromosomal DNA structural heterogeneity from long-read sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.