Novel joint enrichment test demonstrates high performance in simulations and identifies cell-types with enriched expression of inflammatory bowel disease risk loci
Voda, A.-I.; Jostins-Dean, L.
Show abstract
A number of methods have been developed to assess the enrichment of polygenic risk variants - from summary statistics of genome-wide association studies (GWAS) - within specific gene-sets, pathways, or cell-type signatures. The assumptions made by these methods vary, which leads to differences in results and performance across different genetic trait architectures and sample sizes. We devise a novel statistical test that combines independent signals from each of three commonly-used enrichment tests (LDSC, MAGMA & SNPsea) into a single P-value, called the block jackknife GWAS joint enrichment test (GWASJET). Through simulations, we show that this method has comparable or greater power than competing methods across a range of sample sizes and trait architectures. We use our new test in an extensive analysis of the cell-type specific enrichment of genetic risk for inflammatory bowel disease (IBD), including Crohns disease (CD) and ulcerative colitis (UC). Counterintuitively, we find stronger enrichments of IBD risk genes in older gene expression data from bulk immune cell-types than in single-cell data from inflamed patient intestinal samples. We demonstrate that GWASJET removes many seemingly-spurious enriched cell-types identified by other methods, and identifies a core set of immune cells that express IBD risk genes, particularly myeloid cells that have been experimentally stimulated. We also demonstrate that many cell-types are differentially enriched for CD compared to UC risk genes, for example gamma-delta T cells show stronger enrichment for CD than UC risk genes. Author summaryGenetic association studies have discovered a number of DNA variations that are associated with heritable human diseases and traits. One method of investigating the functions of these variants is to test whether they are enriched in parts of the genome associated with specific cell-types or cell conditions - defined by gene expression data or other similar data types. However, there are a number of published statistical methods to test such enrichments; these methdos make different assumptions and their results can vary, sometimes dramatically. We present a novel consensus method, called GWASJET, that combines the results of these different methods to produce a single result. We show that GWASJET can outperform individual methods in simulations. We apply this method to gene expression data from a number of tissues and conditions relevant to inflammatory bowel diseases (IBD). Our method removes potentially false results based on a priori biological knowledge, and reveals that IBD genes are generally clustered in a large number of immune cell-types, especially myeloid cells treated with specific stimulatory molecules.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adjusting for principal components can induce spurious associations in genome-wide association studies in admixed populations 96%
- Leveraging expression from multiple tissues using sparse canonical correlation analysis (sCCA) and aggregate tests improves the power of transcriptome-wide association studies (TWAS) 96%
- Identifying Causal Variants by Fine Mapping Across Multiple Studies 96%
Similar papers in this journal
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 97%
- Sparse modeling of interactions enables fast detection of genome-wide epistasis in biobank-scale studies 96%
- Multiple-testing corrections in selection scans using identity-by-descent segments 96%
Similar papers in this journal
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 97%
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 94%
- Single locus theory of admixture is insufficient for the study of complex traits in admixed populations 94%
Similar papers in this journal
- Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. 96%
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 95%
- Leveraging a machine learning derived surrogate phenotype to improve power for genome-wide association studies of partially missing phenotypes in population biobanks 95%
Similar papers in this journal
- Accurate modeling of replication rates in genome-wide association studies by accounting for winner's curse and study-specific heterogeneity 95%
- Estimating the effective sample size in association studies of quantitative traits 94%
- Restricted maximum-likelihood method for learning latent variance components in gene expression data with known and unknown confounders 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.