Back

Genome assembly with variable order de Bruijn graphs

Diaz, D.; Salmela, L.; Puglisi, S. J.; Onodera, T.

2022-09-07 bioinformatics
10.1101/2022.09.06.506758 bioRxiv
Show abstract

Choosing an order for constructing a de Bruijn graph (DBG) is a crucial step in de novo assembly, as no single value allows complete genome reconstruction. The variable-order de Bruijn graph (voDBG) addresses this limitation by combining DBGs of multiple orders in a single structure connected by contextual relationships. This representation enables new connections to be identified or ambiguities to be resolved during assembly. However, voDBGs currently lack a formal definition of contigs. In this paper, we give the first formal definition of contigs for voDBGs. We show that, for a frequency range [{ell}, h] with{ell} > h/2, nodes whose labels occur with frequency f [isin] [{ell}, h] in the reads spell sequences of the genome with high probability under uniform sampling assumptions. We call these sequences ({ell}, h)-tigs. We also present an efficient algorithm to enumerate ({ell}, h)-tigs from a voDBG that accounts for homopolymer errors. Experiments on PacBio HiFi data show that our method significantly improves contiguity compared to fixed-order DBGs while remaining considerably lighter than full genome assemblers.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.