SNVformer: An Attention-based Deep Neural Network for GWAS Data
Elmes, K.; Benavides-Prado, D.; Tan, N. O.; Nguyen, T. B.; Witbrock, M.; Gavryushkin, A.
Show abstract
Despite being the widely-used gold standard for linking common genetic variations to phenotypes and disease, genome-wide association studies (GWAS) suffer major limitations, partially attributable to the reliance on simple, typically linear, models of genetic effects. More elaborate methods, such as epistasis-aware models, typically struggle with the scale of GWAS data. In this paper, we build on recent advances in neural networks employing Transformer-based architectures to enable such models at a large scale. As a first step towards replacing linear GWAS with a more expressive approximation, we demonstrate prediction of gout, a painful form of inflammatory arthritis arising when monosodium urate crystals form in the joints under high serum urate conditions, from Single Nucleotide Variants (SNVs) using a scalable (long input) variant of the Transformer architecture. Furthermore, we show that sparse SNVs can be efficiently used by these Transformer-based networks without expanding them to a full genome. By appropriately encoding SNVs, we are able to achieve competitive initial performance, with an AUROC of 83% when classifying a balanced test set using genotype and demographic information. Moreover, the confidence with which the network makes its prediction is a good indication of the prediction accuracy. Our results indicate a number of opportunities for extension, enabling full genome-scale data analysis using more complex and accurate genotype-phenotype association models.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PENGUINN: Precise Exploration of Nuclear G-quadruplexes Using Interpretable Neural Networks 94%
- Machine Learning Approaches Identify Genes Containing Spatial Information from Single-Cell Transcriptomics Data. 93%
- Impute.me: an open source, non-profit tool for using data from DTC genetic testing to calculate and interpret polygenic risk scores. 93%
Similar papers in this journal
- LoFTK: a framework for fully automated calculation of predicted Loss-of-Function variants 92%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 92%
- Expanding a Database-derived Biomedical Knowledge Graph via Multi-relation Extraction from Biomedical Abstracts 91%
Similar papers in this journal
- Accurate Prediction of Virus-Host Protein-Protein Interactions via a Siamese Neural Network Using Deep Protein Sequence Embeddings 93%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 93%
- Hierarchical confounder discovery in the experiment-machine learning cycle 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.