It's a wrap: deriving distinct discoveries with FDR control after a GWAS pipeline
Chu, B.; He, Z.; Sabatti, C.
Show abstract
The standard analysis pipeline for genome-wide association studies (GWAS) is based on marginal tests of association. These are computationally convenient and portable, but the discoveries resulting from their rejections are not immediately interpretable, and require post-processing as "clumping" and "fine mapping." An interesting alternative is provided by conditional independence hypotheses: their rejections lead to the identification of distinct signals across the genome, accounting for measured confounders, and pointing to separate causal pathways. An obstacle to the wide adoption of this approach has been that it requires access to individual level data. Overcoming this barrier, recent work has shown how summary statistics resulting from the standard marginal GWAS analysis can be used as input of a procedure to test conditional independence hypotheses while controlling the false discovery rate. This secondary analysis requires sampling of synthetic negative controls (knockoffs) from a distribution determined by the linkage disequilibrium patterns in the genome of the population under study. In prior work, we have pre-computed this distribution for European genomes, starting from information derived from the UK Biobank. Thus, researchers working with GWAS in a European population can carry out a knockoff analysis with minimal computational costs, using the distributed routine GhostKnockoffGWAS. Here we introduce and release a new software (solveblock) that extends this capability to a much richer collection of studies. Given a set of genotyped samples, or a reference dataset, our pipeline efficiently estimates the high-dimensional correlation matrices that describe dependencies across the genome, making rather common sparsity assumptions. Taking this sample-specific estimate as input, the software identifies groups of genetic variants that are highly correlated, and uses them to define an appropriate resolution for conditional independence hypotheses. Finally, we compute the distribution for the exchangeable negative controls necessary to test these hypotheses. The output of solveblock can be passed directly to GhostKnockoffGWAS, allowing users to carry out the complete analysis in a two step procedure. Simulations, based on five UK Biobank sub-populations, illustrate the methods FDR control. The analysis of 26 phenotypes of varying polygenicity in British individuals, results in{approx} 19 additional discoveries, compared to standard marginal association testing. Our code, precompiled software, and processed files for these five sub-populations are openly shared.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying Causal Variants by Fine Mapping Across Multiple Studies 97%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 97%
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 96%
Similar papers in this journal
- Efficient approaches for large scale GWAS studies with genotype uncertainty 95%
- NGSremix: A software tool for estimating pairwise relatedness between admixed individuals from next-generation sequencing data 95%
- gJLS2: A generalized joint location and scale analysis tool for X-inclusive genome-wide discoveries 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.