GWANN: Implementing deep learning in genome wide association studies
Ashkenazy, N.; Feder, M.; Shir, O.; Hubner, S.
Show abstract
MotivationGenome wide association studies (GWAS) are extensively used across species to identify genes that underlie important traits. Most GWAS methods apply modifications and extensions to a linear regression model in order to detect significant associations between genetic variation and a trait. Despite their popularity, these statistical models tend to suffer from high false positive rates, especially when utilized on large variant datasets or complex demographic scenarios. To overcome this, aggressive statistical corrections are applied which frequently diminish true associations. ResultsHere we consider a deep learning approach, and present an implementation of a convolutional neural network (CNN) to identify genetic variation that is associated with a trait of interest. To exploit the strength of CNNs in visual recognition, the genotype information is represented as an image, which enables the model to correctly classify genetic variants with respect to the trait, even when a population structure is present. Our proposed approach was implemented in a package called GWANN which exhibited solid performance. Overall, GWANN outperformed popular GWAS tools on both simulated and real datasets, and enabled the identification of association signals with increased sensitivity and speed. Availability and implementationThe package is available at: https://github.com/hubner-lab/GWANN
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LEA 3: Factor models in population genetics and ecological genomics with R 96%
- Genomic prediction of individual inbreeding levels for the management of genetic diversity in populations with small effective size 95%
- SambaR: an R package for fast, easy and reproducible population-genetic analyses of biallelic SNP datasets 94%
Similar papers in this journal
- A deep learning framework for characterization of genotype data 97%
- Interpretable Artificial Neural Networks incorporating Bayesian Alphabet Models for Genome-wide Prediction and Association Studies 96%
- <underline>Genome-wide Imputation Using the Practical Haplotype Graph in the Heterozygous Crop Cassava</underline> 95%
Similar papers in this journal
- Efficient Permutation-based Genome-wide Association Studies for Normal and Skewed Phenotypic Distributions 94%
- AlphaFamImpute: high accuracy imputation in full-sib families from genotype-by-sequencing data 94%
- The Practical Haplotype Graph, a platform for storing and using pangenomes for imputation 94%
Similar papers in this journal
- Using local convolutional neural networks for genomic prediction 95%
- Strategies to assure optimal trade-offs among competing objectives for genetic improvement of soybean 93%
- Accelerated matrix-vector multiplications for matrices involving genotype covariates with applications in genomic prediction 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.