Back

GnnDebugger: GNN based error correction in De Bruijn Graphs

Simunovic, M.; Sikic, M.; Bankevich, A.

2025-05-13 bioinformatics
10.1101/2025.05.07.652713 bioRxiv
Show abstract

MotivationModern sequencing technologies have enabled the reconstruction of complete mammalian genomes from telomere to telomere. However, scaling this achievement to thousands of species and population-level studies remains a challenge. Key bottlenecks include the low quality of the draft assemblies and the high coverage requirements. In particular, reconstructing complete and accurate sequences of both haplotypes in diploid genomes is especially difficult since the sequencing depth is not always sufficient to properly reconstruct diverged regions. Inspired by the success of neural networks in extracting patterns from the data on a massive scale, we introduce a method for correcting errors in De Bruijn Graphs using Graph Neural Networks. ResultsOur model provides a reliable classification of edges into correct and erroneous, especially for diploid genomes with coverage depth 35 and lower. We demonstrate that these predictions can guide the downstream read error correction algorithm and genome assembly, ultimately allowing for more accurate genome assembly. Availability and implementationBoth GnnDebugger (https://github.com/m5imunovic/gnndebugger) and LJA (https://github.com/AntonBankevich/LJA/tree/gnndebugger) are available on GitHub. Datasets used for training and testing of ML model are available at Zenodo: https://doi.org/10.5281/zenodo.15073168. HG002 reference and reads are available at https://github.com/marbl/HG002. Primates references and reads are available at https://github.com/marbl/Primates.

Published in BMC Bioinformatics (predicted rank #3) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.