Back

PMBB Geno-Pheno Toolkit: A suite of scalable, reproducible pipelines for cross-biobank association analyses

Rodriguez, Z. B.; Guare, L.; Caruth, L.; Cardone, K. M.; Carson, C. C.; Cherlin, T.; Mohammed, S.; Gupta, H.; Kumar, R.; Keat, K.; Verma, S. S.; Verma, A.

2026-07-26 bioinformatics
10.64898/2026.07.22.740077 bioRxiv
Show abstract

SummaryElectronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkits SAIGE family of pipelines -- supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing -- together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted. AvailabilityThe PMBB Geno-Pheno Toolkit is freely available at https://github.com/PMBB-Informatics-and-Genomics/pmbb-geno-pheno-toolkit under MIT open-source license.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.