Automatic Generation of Model Sequences for Complex Regions in Assembly Graphs
Antipov, D.; Chen, Y.; Sollitto, M.; Phillippy, A. M.; Formenti, G.; Koren, S.
Show abstract
Recent developments in genome sequencing and assembly technologies have enabled the automated assembly of vertebrate chromosomes from telomere to telomere. However, for some long, highly similar repeats, genome assemblers may lack sufficient information to unambiguously resolve the sequence, leaving tangles in the assembly graph and gaps in the final assembly. In recently published genomes, such gaps are often closed by manual graph curation, a process that is labor-intensive, error-prone, and sometimes infeasible. This can leave important genomic repeats, such as recently duplicated genes, misassembled or excluded from the final assembly. Here we present the Trivial Tangle Traverser (TTT) algorithm that finds optimized resolutions of assembly graph tangles. TTT uses depth of coverage and read-to-graph alignment information in a two-stage process to identify evidence-based traversals that are consistent with the underlying data. First, sequence multiplicities are estimated through mixed-integer linear programming, after which an Eulerian path is found in the derived multigraph and optimized through a gradient-descent-like approach. We evaluate TTT traversals on the HG002 human reference genome and demonstrate its use to characterize a previously unassembled amplified gene array in the zebra finch genome. AvailabilityTTT is available at https://github.com/marbl/TTT
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A High-quality Oxford Nanopore Assembly of the Hourglass Dolphin (Lagenorhynchus cruciger) Genome 95%
- A new high-quality genome assembly and annotation for the threatened Florida Scrub-Jay (Aphelocoma coerulescens) 95%
- PAQman: reference-free ensemble evaluation of long-read eukaryotic genome assemblies 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.