AsmMix: A pipeline for high quality diploid denovo assembly
Wu, P.; Liu, C.; Wang, O.; Xia, Z.; Chen, F.; Chen, X.; Zhu, H.
Show abstract
In this paper, we report a pipeline, AsmMix, which is capable of producing both contiguous and high-quality diploid genomes. The pipeline consists of two steps. In the first step, two sets of assemblies are generated: one is based on co-barcoded reads, which are highly accurate and haplotype-resolved but contain many gaps, the other assembly is based on single-molecule sequencing reads, which is contiguous but error-prone. In the second step, those two sets of assemblies are compared and integrated into a haplotype-resolved assembly with fewer errors. We test our pipeline using a dataset of human genome NA24385, perform variant calling from those assemblies and then compare against GIAB Benchmark. We show that AsmMix pipeline could produce highly contiguous, accurate, and haplotype-resolved assemblies. Especially the assembly mixing process could effectively reduce small-scale errors in the long read assembly.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaBinner: a high-performance and stand-alone ensemble binning method to recover individual genomes from complex microbial communities 96%
- deSALT: fast and accurate long transcriptomic read alignment with de Bruijn graph-based index 96%
- Long-read-based Human Genomic Structural Variation Detection with cuteSV 96%
Similar papers in this journal
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 94%
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 94%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 94%
Similar papers in this journal
- Real-time resolution of short-read assembly graph using ONT long reads 97%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 95%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.