Application of post-selection inference to multi-omics data yields insights into the etiologies of human diseases
Yurko, R. J.; G'Sell, M.; Roeder, K.; Devlin, B.
Show abstract
To correct for a large number of hypothesis tests, most researchers rely on simple multiple testing corrections. Yet, new methodologies of selective inference could potentially improve power while retaining statistical guarantees, especially those that enable exploration of test statistics using auxiliary information (covariates) to weight hypothesis tests for association. We explore one such method, adaptive p-value thresholding (Lei & Fithian 2018, AdaPT), in the framework of genome-wide association studies (GWAS) and gene expression/coexpression studies, with particular emphasis on schizophrenia (SCZ). Selected SCZ GWAS association p-values play the role of the primary data for AdaPT; SNPs are selected because they are gene expression quantitative trait loci (eQTLs). This natural pairing of SNPs and genes allow us to map the following covariate values to these pairs: GWAS statistics from genetically-correlated bipolar disorder, the effect size of SNP genotypes on gene expression, and gene-gene coexpression, captured by subnetwork (module) membership. In all 24 covariates per SNP/gene pair were included in the AdaPT analysis using flexible gradient boosted trees. We demonstrate a substantial increase in power to detect SCZ associations using gene expression information from the developing human prefontal cortex (Werling et al. 2019). We interpret these results in light of recent theories about the polygenic nature of SCZ. Importantly, our entire process for identifying enrichment and creating features with independent complementary data sources can be implemented in many different high-throughput settings to ultimately improve power.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RapidoPGS: A rapid polygenic score calculator for summary GWAS data without a test dataset 96%
- Nearest-neighbor Projected-Distance Regression (NPDR) for detecting network interactions with adjustments for multiple tests and confounding 96%
- Sparse Polygenic Risk Score Inference with the Spike-and-Slab LASSO 95%
Similar papers in this journal
- Efficient gene-environment interaction tests for large biobank-scale sequencing studies 96%
- RetroFun-RVS: a retrospective family-based framework for rare-variant analysis incorporating functional annotations 95%
- Identity-by-descent mapping using multi-individual IBD with genome-wide multiple testing adjustment 94%
Similar papers in this journal
Similar papers in this journal
- Estimating the effective sample size in association studies of quantitative traits 94%
- gJLS2: A generalized joint location and scale analysis tool for X-inclusive genome-wide discoveries 93%
- Restricted maximum-likelihood method for learning latent variance components in gene expression data with known and unknown confounders 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.