Back

Cloud gazing: demonstrating paths for unlocking the value of cloud genomics through cross-cohort analysis

Deflaux, N.; Selvaraj, M. S.; Condon, H. R.; Mayo, K.; Haidermota, S.; Basford, M. A.; Lunt, C.; Philippakis, A. A.; Roden, D. M.; Denny, J. C.; Musick, A.; Collins, R.; Allen, N.; Effingham, M.; Glazer, D.; Natarajan, P.; Bick, A. G.

2022-12-02 genetics
10.1101/2022.11.29.518423 bioRxiv
Show abstract

The rapid growth of genomic data has led to a new research paradigm where data are stored centrally in Trusted Research Environments (TREs) such as the All of Us Researcher Workbench (AoU RW) and the UK Biobank Research Analysis Platform (RAP). To characterize the advantages and drawbacks of different TRE attributes in facilitating cross-cohort analysis, we conducted a Genome-Wide Association Study (GWAS) of standard lipid measures on the UKB RAP and AoU RW using two approaches: meta-analysis and pooled analysis. We curated lipid measurements for 37,754 All of Us participants with whole genome sequence (WGS) data and 190,982 UK Biobank participants with whole exome sequence (WES) data. For the meta-analysis, we performed a GWAS of each cohort in their respective platform and meta-analyzed the results. We separately performed a pooled GWAS on both datasets combined. We identified 490 and 464 significant variants in meta-analysis and pooled analysis, respectively. Comparison of full summary data from both meta-analysis and pooled analysis with an external study showed strong correlation of known loci with lipid levels (R2[~]83-97%). Importantly, 90 variants met the significance threshold only in the meta-analysis and 64 variants were significant only in pooled analysis. These method-specific differences may be explained by differences in cohort size, ancestry, and phenotype distributions in All of Us and UK Biobank. We noted approximately 20% of variants significant in only the pooled analysis or significant in only the meta-analysis were most prevalent in non-European, non-Asian ancestry individuals. Pooled analyses included more variants than meta-analyses. Pooled analysis required about half as many computational steps as meta-analysis. These findings have important implications for both platform implementations and researchers undertaking large-scale cross-cohort analyses, as technical and policy choices lead to cross-cohort analyses generating similar, but not identical results, particularly for non-European ancestral populations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.