PennPRS: a centralized cloud computing platform for efficient polygenic risk score training in precision medicine
Jin, J.; Li, B.; Wang, X.; Yang, X.; Li, Y.; Wang, R.; Ye, C.; Shu, J.; Fan, Z.; Xue, F.; Ge, T.; Ritchie, M. D.; Pasaniuc, B.; Wojcik, G.; Zhao, B.
Show abstract
Polygenic risk scores (PRS) are becoming increasingly vital for risk prediction and stratification in precision medicine. However, PRS model training presents significant challenges for broader adoption of PRS, including limited access to computational resources, difficulties in implementing advanced PRS methods, and availability and privacy concerns over individual-level genetic data. Cloud computing provides a promising solution with centralized computing and data resources. Here we introduce PennPRS (https://pennprs.org), a scalable cloud computing platform for online PRS model training in precision medicine. We developed novel pseudo-training algorithms for multiple PRS methods and ensemble approaches, enabling model training without requiring individual-level data. These methods were rigorously validated through extensive simulations and large-scale real data analyses involving over 6,000 phenotypes across various data sources. PennPRS supports online single- and multi-ancestry PRS training with seven methods, allowing users to upload their own data or query from more than 27,000 datasets in the GWAS Catalog, submit jobs, and download trained PRS models. Additionally, we applied our pseudo-training pipeline to train PRS models for over 8,000 phenotypes and made their PRS weights publicly accessible. In summary, PennPRS provides a novel cloud computing solution to improve the accessibility of PRS applications and reduce disparities in computational resources for the global PRS research community.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A new method for multi-ancestry polygenic prediction improves performance across diverse populations 97%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 97%
- Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies 97%
Similar papers in this journal
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 96%
- MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups 96%
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 95%
Similar papers in this journal
- OTTERS: A powerful TWAS framework leveraging summary-level reference data 97%
- Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics 97%
- Uncovering causal gene-tissue pairs and variants: A multivariable TWAS method controlling for infinitesimal effects 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.