Back

Archipelago method for variant set association test statistics.

Lawless, D.; Saadat, A.; Oumelloul, M. A.; Fellay, J.

2025-03-17 genetic and genomic medicine
10.1101/2025.03.17.25324111 medRxiv
Show abstract

Variant set association tests (VSAT), especially those incorporating rare variants via variant collapse, are invaluable in genetic studies. However, unlike Manhattan plots for single-variant tests, VSAT statistics lack intrinsic genomic coordinates, hindering visual interpretation. To overcome this, we developed the Archipelago method, which assigns a meaningful genomic coordinate to VSAT P values so that both set-level and individual variant associations can be visualised together. This results in an intuitive and information rich illustration akin to an Archipelago of clustered islands, enhancing the understanding of both collective and individual impacts of variants. We conducted three validation studies spanning simulated and real datasets across small and biobank-scale cohorts, from 504 individuals up to 490,640 UK Biobank participants. We integrated single-variant genome-wide association studies (GWAS) with gene-and protein pathway-level rare-variant collapse. These studies included the 1KG GWAS cohort, the Pan-UK Biobank GWAS with DeepRVAT WES gene-level study, and the UKBB WGS gene-level UTR collapsing PheWAS. The Archipelago plot is applicable in any genetic association study that uses variant collapse to evaluate both individual variants and variant sets, and its customisability facilitates clear communication of complex genetic data. By integrating at least two dimensions of genetic data into a single visualisation, VSAT results can be easily read and aid in identification of potential causal variants in variant sets such as protein pathways. GitHub repositoryhttps://github.com/DylanLawless/archipelago. CRAN submissionArchipelago Version 0.0.1.9000. Zenodohttps://doi.org/10.5281/zenodo.16880622.

Published in Genetic Epidemiology (predicted rank #2) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.