Back

Investigating the topological motifs of inversions in pangenome graphs

Romain, S.; Dubois, S.; Legeai, F.; Lemaitre, C.

2025-03-17 bioinformatics
10.1101/2025.03.14.643331 bioRxiv
Show abstract

BackgroundPangenome graphs are increasingly used in genetic diversity anal-yses because they reduce reference bias in read mapping and enhance variant discovery and genotyping from SNPs to Structural Variants. In pangenome graphs, variants appear as bubbles, which can be detected by dedicated bubble calling tools. Although these tools report essential information on the variant bubbles, such as their position and allele walks in the graph, they do not annotate the type of the detected variants. While simple SNPs, insertions, and deletions are easily distinguishable by allele size, large balanced variants like inversions are harder to differentiate among the large number of unannotated bubbles and remain underexplored in pangenome graph benchmarks and analyses. ResultsIn this work we focused on inversions, which have been drawing renewed attention in evolutionary genomics studies in the past years, and aimed to assess how this type of variant is handled by state of the art pangenome graph pipelines. We identified two distinct topological motifs for inversion bubbles: one path-explicit and one alignment-rescued, and developed a tool to annotate them from bubble-caller outputs. We constructed pangenome graphs with both simulated data and real data using four state of the art pipelines, and assessed the impact of inversion size, genome divergence and variant density on inversion representation and accuracy. ConclusionsOur results reveal substantial differences between pipelines in simulated graphs, with some inversions either misrepresented or lost. In addition, recovery rates are strikingly low in real human datasets, highlighting major challenges in analyzing inversions through pangenomic approaches.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.