GPatch enables chromosome-scale, gap-free pseudoassemblies from fragmented draft genomes
Diehl, A.; Boyle, A. P.
Show abstract
Recent advancements in sequencing technologies have yielded numerous long-read draft genomes, promising to enhance understanding of genomic variation. However, draft genomes are typically highly fragmented, posing challenges for functional genomics. We introduce GPatch, a tool that constructs chromosome-scale pseudoassemblies from fragmented drafts using alignments to a reference assembly. GPatch produces complete, accurate, gap-free assemblies preserving over 95% of nucleotides from human and non-human draft genomes. We show that GPatch pseudoassemblies can be used to construct Hi-C matrices, whereas fragmented draft assemblies cannot. Until complete genome assembly becomes routine, GPatch presents a necessary tool for maximizing the utility of draft genomes. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/655567v2_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@5e3a41org.highwire.dtl.DTLVardef@42bf23org.highwire.dtl.DTLVardef@12cd6forg.highwire.dtl.DTLVardef@6d720c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Gaps and complex structurally variant loci in phased genome assemblies 96%
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 96%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
Similar papers in this journal
- Pushing the limits of HiFi assemblies reveals centromere diversity between two Arabidopsis thaliana genomes 96%
- sgcocaller and comapr: personalised haplotype assembly and comparative crossover map analysis using single-gamete sequencing data 95%
- Dysgu: efficient structural variant calling using short or long reads 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.