Multitask group Lasso for Genome Wide association Studies in admixed populations
Nouira, A.; Azencott, C.-A.
Show abstract
Genome-Wide Association Studies, or GWAS, aim at finding Single Nucleotide Polymorphisms (SNPs) that are associated with a phenotype of interest. GWAS are known to suffer from the large dimensionality of the data with respect to the number of available samples. Other limiting factors include the dependency between SNPs, due to linkage disequilibrium (LD), and the need to account for population structure, that is to say, confounding due to genetic ancestry. We propose an efficient approach for the multivariate analysis of multi-population GWAS data based on a multitask group Lasso formulation. Each task corresponds to a subpopulation of the data, and each group to an LD-block. This formulation alleviates the curse of dimensionality, and makes it possible to identify disease LD-blocks shared across populations/tasks, as well as some that are specific to one population/task. In addition, we use stability selection to increase the robustness of our approach. Finally, gap safe screening rules speed up computations enough that our method can run at a genome-wide scale. To our knowledge, this is the first framework for GWAS on diverse populations combining feature selection at the LD-groups level, a multitask approach to address population structure, stability selection, and safe screening rules. We show that our approach outperforms state-of-the-art methods on both a simulated and a real-world cancer datasets.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BEATRICE: Bayesian Fine-mapping from Summary Datausing Deep Variational Inference 96%
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 96%
- Sparse Polygenic Risk Score Inference with the Spike-and-Slab LASSO 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 97%
- Leveraging a machine learning derived surrogate phenotype to improve power for genome-wide association studies of partially missing phenotypes in population biobanks 97%
- Computationally efficient whole genome regression for quantitative and binary traits 97%
Similar papers in this journal
- Identification of putative causal loci in whole-genome sequencing data via knockoff statistics 96%
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 96%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.