Genome assembly with variable order de Bruijn graphs
Diaz, D.; Salmela, L.; Puglisi, S. J.; Onodera, T.
Show abstract
Choosing an order for constructing a de Bruijn graph (DBG) is a crucial step in de novo assembly, as no single value allows complete genome reconstruction. The variable-order de Bruijn graph (voDBG) addresses this limitation by combining DBGs of multiple orders in a single structure connected by contextual relationships. This representation enables new connections to be identified or ambiguities to be resolved during assembly. However, voDBGs currently lack a formal definition of contigs. In this paper, we give the first formal definition of contigs for voDBGs. We show that, for a frequency range [{ell}, h] with{ell} > h/2, nodes whose labels occur with frequency f [isin] [{ell}, h] in the reads spell sequences of the genome with high probability under uniform sampling assumptions. We call these sequences ({ell}, h)-tigs. We also present an efficient algorithm to enumerate ({ell}, h)-tigs from a voDBG that accounts for homopolymer errors. Experiments on PacBio HiFi data show that our method significantly improves contiguity compared to fixed-order DBGs while remaining considerably lighter than full genome assemblers.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.