Back

Defining and cataloging variants in pangenome graphs

Salehi Nowbandegani, P.; Zhang, S.; Hu, H.; Li, H.; O'Connor, L. J.

2025-08-04 genomics
10.1101/2025.08.04.668502 bioRxiv
Show abstract

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to reference bias. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a genetic variant. We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Published in Cell Genomics (predicted rank #9) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.