Variant Set Distillation
Christ, R.; Kang, C. J.; Aslett, L.; Lam, D.; Savitski, M. F.; Stitziel, N. O.; Steinsaltz, D.; Hall, I.
Show abstract
Allelic heterogeneity - the presence of multiple causal variants at a given locus - has been widely observed across human traits. Combining the association signals across these distinct causal variants at a given locus presents an opportunity for empowering gene discovery. This opportunity is growing with the increasing population diversity and sequencing depth of emerging genomic datasets. However, the rapidly increasing number of null (non-causal) variants within these datasets makes leveraging allelic heterogeneity increasingly difficult for existing testing approaches. We recently-proposed a general theoretical framework for sparse signal problems, Stable Distillation (SD). Here we present a SD-based method vsdistill, which overcomes several major shortcomings in the simple SD procedures we initially proposed and introduces many innovations aimed at maximizing power in the context of genomics. We show via simulations that vsdistill provides a significant power boost over the popular STAAR method. vsdistill is available in our new R package gdistill, with core routines implemented in C. We also show our method scales readily to large datasets by performing an association analysis with height in the UK Biobank.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 96%
- Enhancing sensitivity and controlling false discovery rate in somatic indel discovery 96%
- MEDICC2: whole-genome doubling aware copy-number phylogenies for cancer evolution 95%
Similar papers in this journal
- ML-MAGES: A machine learning framework for multivariate genetic association analyses with genes and effect size shrinkage 97%
- Leveraging family data to design Mendelian Randomization that is provably robust to population stratification 97%
- Assessing transcriptomic re-identification risks using discriminative sequence models 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.