Back

isGWAS: ultra-high-throughput, scalable and equitable inference of genetic associations with disease

Foley, C. N.; Kuncheva, Z.; Marioni, R.; Runz, H.; Sun, B.

2023-08-10 genetics
10.1101/2023.07.21.550074 bioRxiv
Show abstract

Genome-wide association studies (GWAS) have proven a powerful tool for human geneticists to generate biological insights or hypotheses for drug discovery. Nevertheless, a dependency on sensitive individual-level data together with ever-increasing cohort sample sizes, numbers of variants and phenotypes studied put a strain on existing algorithms, limiting the GWAS approach from maximising potential. Here we present in-silico GWAS (isGWAS), a uniquely scalable algorithm to infer regression parameters in case-control GWAS from cohort-level summary data. For any sample size, isGWAS computes a variant-disease association parameter in [~]1 millisecond, or [~]11m variants in UK-Biobank within [~]4 minutes ([~]1500-fold faster than state-of-the-art). Extensive simulations and empirical tests demonstrate that isGWAS results are highly comparable to traditional regression-based approaches. We further introduce a heuristic re-sampling algorithm, leapfrog re-sampler (LRS), to extrapolate association results to semi-virtually enlarged cohorts. Owing to significant computational gains we anticipate a broad use of isGWAS and LRS which are customizable on a web interface.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.