Cloud gazing: demonstrating paths for unlocking the value of cloud genomics through cross-cohort analysis
Deflaux, N.; Selvaraj, M. S.; Condon, H. R.; Mayo, K.; Haidermota, S.; Basford, M. A.; Lunt, C.; Philippakis, A. A.; Roden, D. M.; Denny, J. C.; Musick, A.; Collins, R.; Allen, N.; Effingham, M.; Glazer, D.; Natarajan, P.; Bick, A. G.
Show abstract
The rapid growth of genomic data has led to a new research paradigm where data are stored centrally in Trusted Research Environments (TREs) such as the All of Us Researcher Workbench (AoU RW) and the UK Biobank Research Analysis Platform (RAP). To characterize the advantages and drawbacks of different TRE attributes in facilitating cross-cohort analysis, we conducted a Genome-Wide Association Study (GWAS) of standard lipid measures on the UKB RAP and AoU RW using two approaches: meta-analysis and pooled analysis. We curated lipid measurements for 37,754 All of Us participants with whole genome sequence (WGS) data and 190,982 UK Biobank participants with whole exome sequence (WES) data. For the meta-analysis, we performed a GWAS of each cohort in their respective platform and meta-analyzed the results. We separately performed a pooled GWAS on both datasets combined. We identified 490 and 464 significant variants in meta-analysis and pooled analysis, respectively. Comparison of full summary data from both meta-analysis and pooled analysis with an external study showed strong correlation of known loci with lipid levels (R2[~]83-97%). Importantly, 90 variants met the significance threshold only in the meta-analysis and 64 variants were significant only in pooled analysis. These method-specific differences may be explained by differences in cohort size, ancestry, and phenotype distributions in All of Us and UK Biobank. We noted approximately 20% of variants significant in only the pooled analysis or significant in only the meta-analysis were most prevalent in non-European, non-Asian ancestry individuals. Pooled analyses included more variants than meta-analyses. Pooled analysis required about half as many computational steps as meta-analysis. These findings have important implications for both platform implementations and researchers undertaking large-scale cross-cohort analyses, as technical and policy choices lead to cross-cohort analyses generating similar, but not identical results, particularly for non-European ancestral populations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Inclusion of Variants Discovered from Diverse Populations Improves Polygenic Risk Score Transferability 95%
- A reference panel for linkage disequilibrium and genotype imputation using whole-genome sequencing data from 2,680 participants across India 94%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 94%
Similar papers in this journal
- Streamlining Large-Scale Genomic Data Management: Insights from the UK Biobank Whole-Genome Sequencing Data 94%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 94%
- Genome-wide study on 72,298 Korean individuals in Korean biobank data for 76 traits identifies hundreds of novel loci 94%
Similar papers in this journal
- A novel Mendelian randomization method identifies causal relationships between gene expression and low-density lipoprotein cholesterol levels. 94%
- Evaluating transportability of in-vitro cellular models to in-vivo human phenotypes using gene perturbation data 94%
- Whole genome sequence analysis of blood lipid levels in >66,000 individuals 94%
Similar papers in this journal
- High-throughput multivariable Mendelian randomization analysis prioritizes apolipoprotein B as key lipid risk factor for coronary artery disease 93%
- An empirical investigation into the impact of winner's curse on estimates from Mendelian randomization 92%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.