Back

On Deriving Synteny Blocks by Compacting Elements

Bohnenkaemper, L.; Parmigiani, L.; Chauve, C.; Stoye, J.

2025-12-17 bioinformatics
10.64898/2025.12.15.694404 bioRxiv
Show abstract

Genomic rearrangements are major drivers of evolution and genetic disease. However, studying rearrangements requires segmenting the genomes of interest into conserved regions, called synteny blocks, that highlight structural differences between genomes. Synteny blocks are typically defined from annotated genes or derived as a by-product of whole-genome alignments. As these procedures are heuristic and do not explicitly model rearrangements, they can obscure real variation, create false similarities, and affect phylogenetic inference. The importance of synteny block definition has long been recognized, as shown for example by discussions on breakpoint reuse, where different definitions of synteny blocks led to different estimates of rearrangement complexity in mammalian genomes. We present a formal framework for deriving synteny blocks directly from sequence data by partitioning genomic elements into blocks that do not contain breakpoints. A breakpoint is defined between a pair of genomes as an adjacency of shared elements that occurs in one genome but not in the other. Synteny blocks are therefore not allowed to span such boundaries, ensuring that rearrangements are not obscured. The framework is fully agnostic to the type of genomic element and applies to any genome representation expressed as sequences of elements, such as non-overlapping alignments, exact matches (MUMs/MEMs), k-mers, unitigs or minimizers. We formalize two optimization problems: minimizing the total genome length after replacement by synteny blocks (the Minimum-Length Synteny Block Problem) and minimizing the number of distinct blocks (the Minimum-Size Synteny Block Problem). We show that both problems are NP-hard in general. However, when blocks are required to be collinear and to contain a shared element, we provide a linear-time algorithm with respect to the number of input elements that simultaneously minimizes both objectives. The resulting method is simple, efficient, and produces large synteny blocks without obscuring rearrangements.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.