PMBB Geno-Pheno Toolkit: A suite of scalable, reproducible pipelines for cross-biobank association analyses
Rodriguez, Z. B.; Guare, L.; Caruth, L.; Cardone, K. M.; Carson, C. C.; Cherlin, T.; Mohammed, S.; Gupta, H.; Kumar, R.; Keat, K.; Verma, S. S.; Verma, A.
Show abstract
SummaryElectronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkits SAIGE family of pipelines -- supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing -- together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted. AvailabilityThe PMBB Geno-Pheno Toolkit is freely available at https://github.com/PMBB-Informatics-and-Genomics/pmbb-geno-pheno-toolkit under MIT open-source license.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 94%
- Impute.me: an open source, non-profit tool for using data from DTC genetic testing to calculate and interpret polygenic risk scores. 94%
- Pipeliner: A Nextflow-based framework for the definition of sequencing data processing pipelines 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.