Back

Anchors for Homology-Based Scaffolding

Kaether, K.-K.; Lemke, S.; Stadler, P. F.

2025-09-13 bioinformatics
10.1101/2025.04.28.650980 bioRxiv
Show abstract

Homology-based scaffolding orders contigs based on conserved collinearity of homologous sequences across related species. Existing methods often rely on costly whole-genome alignments or show limited robustness when integrating multiple references. Here, we introduce an anchor-based scaffolding framework that adapts synteny anchors to efficiently infer contig order and orientation relative to one or more reference genomes. Our approach leverages precomputed, sufficiently unique anchors and their respective high-confidence homology matches in a greedy approach, combining single-reference to multi-reference scaffolds using a maximum matching. Across simulated and real datasets, anchor-based scaffolding achieves accuracy comparable to state-of-the-art methods. Notably, the approach shows particular strengths in multi-reference settings. These results demonstrate that synteny-anchor-based scaffolding provides an additional tool for homology-based scaffolding with robust accuracy and superior performance in multi-reference scenarios.

Published in Journal of Bioinformatics and Computational Biology · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.