GWASHub: An Automated Cloud-Based Platform for Genome-Wide Association Study Meta-Analysis
Sunderland, N.; Hite, D.; Smadbeck, P.; Hoang, Q.; Jang, D.-K.; Tragante, V.; Jiang, J. C.; Shah, S.; Paternoster, L.; Burtt, N. P.; Flannick, J.; Lumbers, R. T.
Show abstract
Genome-wide association studies (GWAS) often aggregate data from millions of participants across multiple cohorts using meta-analysis to maximise power for genetic discovery. The increase in availability of genomic biobanks, together with a growing focus on phenotypic subgroups, genetic diversity, and sex-stratified analyses, has led GWAS meta-analyses to routinely produce hundreds of summary statistic files accompanied by detailed meta-data. Scalable infrastructures for data handling, quality control (QC), and meta-analysis workflows are essential to prevent errors, ensure reproducibility, and reduce the burden on researchers, allowing them to focus on downstream research and clinical translation. To address this need, we developed GWASHub, a secure cloud-based platform designed for the curation, processing and meta-analysis of GWAS summary statistics. GWASHub features i) private and secure project spaces, ii) automated file harmonisation and data validation, iii) GWAS meta-data capture, iv) customisable variant QC, v) GWAS meta-analysis, vi) analysis reporting and visualisation, and vii) results download. Users interact with the portal via an intuitive web interface built on Nuxt.js, a high-performance JavaScript framework. Data is securely managed through an Amazon Web Services (AWS) MySQL database and S3 block storage. Analysis jobs are distributed to AWS compute resources in a scalable fashion. The QC dashboard presents tabular and graphical QC outputs allowing manual review of individual datasets. Those passing QC are made available to the meta-analysis module. Individual datasets and meta-analysis results are available for download by project users with appropriate access permissions. In GWASHub, a "project" serves as a virtual workspace spanning an entire consortium, allowing individuals with different roles, such as data contributors (users) and project coordinators (main analysts), to collaborate securely under a unified framework. GWASHub has a flexible architecture to allow for ongoing development and incorporation of alternative quality control or meta-analysis procedures, to meet the specific needs of researchers. GWASHub was developed as a joint initiative by the HERMES Consortium and the Cardiovascular Knowledge Portal, and access to the platform is free and available upon request. GWASHub addresses a critical need in the genetics research community by providing a scalable, secure, and user-friendly platform for managing the complexity of large-scale GWAS meta-analyses. As the volume and diversity of GWAS data continue to grow, platforms like GWASHub may help to accelerate insights into the genetic architecture of complex traits.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- UK BioCoin: Swift Trait-Specific Summary Statistics Regression for UK Biobank 94%
- Orchestrating and sharing large multimodal data for transparent and reproducible research 93%
- A user's guide to the online resources for data exploration, visualization, and discovery for the Pan-Cancer Analysis of Whole Genomes project (PCAWG) 93%
Similar papers in this journal
- Performing highly parallelized and reproducible GWAS analysis on biobank-scale data 97%
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 94%
- DoBSeqWF: A framework for sensitive detection of individual genetic variation in pooled sequencing data 94%
Similar papers in this journal
- Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches 93%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 92%
- Validation of a Trans-Ancestry Polygenic Risk Score for Type 2 Diabetes in Diverse Populations 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.