Accelerating Genome- and Phenome-Wide Association Studies using GPUs - A case study using data from the Million Veteran Program
Rodriguez, A. A.; Kim, Y.; Nandi, T. N.; Keat, K.; Bhukar, R.; Conery, M.; Liu, M.; Hessington, J.; Begoli, E.; Tourassi, G.; Muralidhar, S.; Natarajan, P.; Voight, B. F.; Cho, K.; Gaziano, M. J.; Damrauer, S.; Liao, K. P.; Zhou, W.; Huffman, J. E.; Verma, A.; Madduri, R. K.
Show abstract
The expansion of biobanks has significantly propelled genomic discoveries yet the sheer scale of data within these repositories poses formidable computational hurdles, particularly in handling extensive matrix operations required by prevailing statistical frameworks. In this work, we introduce computational optimizations to the SAIGE (Scalable and Accurate Implementation of Generalized Mixed Model) algorithm, notably employing a GPU-based distributed computing approach to tackle these challenges. We applied these optimizations to conduct a large-scale genome-wide association study (GWAS) across 2,068 phenotypes derived from electronic health records of 635,969 diverse participants from the Veterans Affairs (VA) Million Veteran Program (MVP). Our strategies enabled scaling up the analysis to over 6,000 nodes on the Department of Energy (DOE) Oak Ridge Leadership Computing Facility (OLCF) Summit High-Performance Computer (HPC), resulting in a 20-fold acceleration compared to the baseline model. We also provide a Docker container with our optimizations that was successfully used on multiple cloud infrastructures on UK Biobank and All of Us datasets where we showed significant time and cost benefits over the baseline SAIGE model.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Torch-eCpG: A fast and scalable eQTM mapper for thousands of molecular phenotypes with graphical processing units 94%
- GeneSetCluster 2.0: a comprehensive toolset for summarizing and integrating gene-sets analysis 93%
- Hypercluster: a flexible tool for parallelized unsupervised clustering optimization 93%
Similar papers in this journal
- CCC-GPU: A graphics processing unit (GPU)-accelerated nonlinear correlation coefficient for large-scale transcriptomic analyses 95%
- Bigtools: a high-performance BigWig and BigBed library in Rust 94%
- MungeSumstats: A Bioconductor package for the standardisation and quality control of many GWAS summary statistics 93%
Similar papers in this journal
- Accelerated matrix-vector multiplications for matrices involving genotype covariates with applications in genomic prediction 92%
- Pipeliner: A Nextflow-based framework for the definition of sequencing data processing pipelines 91%
- SSAM-lite: a light-weight web app for rapid analysis of spatially resolved transcriptomics data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.