Back

Pangenome Graph Node-Phenotype Association shows GWAS-like quality results with only few individuals

Carrette, C.; Sabot, F.; Muller, C.

2026-08-01 bioinformatics
10.64898/2026.07.31.741971 bioRxiv
Show abstract

PurposeWe introduce GO_SCPLOWRAC_SCPLOWNPA, standing for Graph Node-Phenotype Association, a method performing a GWAS-like analysis on a pangenome variation graph (PVG) built using a small number of individual genome sequences, without the need for additional population materials or kinship information for qualitative phenotypes. This method reduces the number of individuals required for association studies and prevents reference bias from variant calling in these types of analyses. BackgroundA PVG represents the multiple alignment of a set of complete genomes. It contains all variations, from single nucleotide polymorphisms (SNPs) to large structural variations (SVs), which are represented as nodes in the graph. By integrating phenotype information within nodes, we can assign a Phenotype Score (PS) to each node in the PVG and identify phenotype-related regions directly within it. These regions represent statistically significant shifts in PS distribution, highlighting their implication in the phenotype. Finally, GO_SCPLOWRAC_SCPLOWNPA provides their positions and scores for further analysis. ResultsThis method was tested using simulated data and two publicly available datasets: the Sub1A gene locus for Oryza sativa in a 13 individuals PVG, and the insertion responsible for the white-headed cattle with a PVG of 24 individuals. Source code of GO_SCPLOWRAC_SCPLOWNPA is available here https://forge.ird.fr/diade/graphgwas/granpa under GNU GPLv3. ConclusionGO_SCPLOWRAC_SCPLOWNPA was able to identify the expected area in two simulated datasets and the responsible loci for these two known traits using only a few dozen complete genomes in these PVGs. While currently limited to qualitative phenotypes, this method opens the way to more efficient ones relying on PVGs and few individuals.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.