GraphUnzip: unzipping assembly graphs with long reads and Hi-C
Faure, R.; Guiglielmoni, N.; Flot, J.-F.
Show abstract
Long reads and Hi-C have revolutionized the field of genome assembly as they have made highly continuous assemblies accessible for challenging genomes. As haploid chromosome-level assemblies are now commonly achieved for all types of organisms, phasing assemblies has become the new frontier for genome reconstruction. Several tools have already been released using long reads and/or Hi-C to phase assemblies, but they all start from a linear sequence, and are ill-suited for non-model organisms with high levels of heterozygosity. We present GraphUnzip, a fast, memory-efficient and accurate tool to unzip assembly graphs into their constituent haplotypes using long reads and/or Hi-C data. As GraphUnzip only connects sequences in the assembly graph that already had a potential link based on overlaps, it yields high-quality gap-less supercontigs. To demonstrate the efficiency of GraphUnzip, we tested it on a simulated diploid Escherichia coli genome, and on two real datasets for the genomes of the rotifer Adineta vaga and the potato Solanum tuberosum. In all cases, GraphUnzip yielded highly continuous phased assemblies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromosome-level quality scaffolding of brown algal genomes using InstaGRAAL, a proximity ligation-based scaffolder 97%
- SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies 97%
- Automated assembly scaffolding elevates a new tomato system for high-throughput genome editing 96%
Similar papers in this journal
- An improved chromosome-level genome assembly of perennial ryegrass (Lolium perenne L.) 95%
- Optimizing experimental design for genome sequencing and assembly with Oxford Nanopore Technologies 95%
- Chromosome-scale assembly of the highly heterozygous genome of red clover (Trifolium pratense L.), an allogamous forage crop species 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.