Back

PolyGenie: An Interactive Platform for Visualising Polygenic Risk Across Multidimensional Cohorts

Farre, X.; Gasco, M.; Blay, N.; de Cid, R.

2025-09-10 genomics
10.1101/2025.09.05.674501 bioRxiv
Show abstract

SummaryPhenome-wide association studies (PheWAS) using polygenic risk scores (PRS) offer a powerful framework for exploring the shared genetic architecture of complex traits across diverse phenotypic domains. However, no standardized, portable pipeline exists to facilitate their systematic execution and visualization in arbitrary population cohorts. We present PolyGenie, an open-source Nextflow pipeline that takes precomputed PRS and cohort phenotype data as input and performs scalable PheWAS analysis across binary and continuous outcomes. The pipeline produces regression results and percentile-based prevalence estimates, which are stored in a SQLite database and visualized through an interactive Dash web application. To demonstrate its utility and reproducibility, we provide a fully worked example using the GCAT cohort, applying 135 PRS to a broad range of clinical, molecular, and lifestyle phenotypes. PolyGenie is designed to be deployed on any cohort with minimal configuration, enabling standardized cross-trait analyses and interactive exploration of genetic risk. Availability and ImplementationPolyGenie is freely available at https://github.com/gcatbiobank/polygenie-pipeline. The GCAT implementation can be accessed at https://polygenie.igtp.cat, and the corresponding adapted Dash source code is available at https://github.com/gcatbiobank/polygenie-gcat. Contactxfarrer@igtp.cat; rdecid@igtp.cat Supplementary informationSupplementary data are available at Bioinformatics online.

Published in NAR Genomics and Bioinformatics · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.