Accuracy and Scalability of Machine Learning Methods for Genotype-Phenotype Association Data
Collienne, K.; Zhang, L.; Gavryushkin, A.
Show abstract
Many machine learning methods can be applied to predicting phenotypes from genetic data. Which of these methods work best remains an open question, however. To answer this question, we propose to compare a variety of approaches ability to predict a simulated non-linear complex trait. Specifically, we evaluate these methods on their accuracy and scalability with respect to the amount training data available, the noise present in the data, the complexity of the simulated (trait) functions, and their ability to provide insight into the simulated trait. We then compare the best approach to state-of-the-art models in real data, predicting gout in the UK Biobank. We find that transformer encoders outperform all other methods in simulations, and perform comparably to the state-of-the-art with real data, with a promise to scale to significantly larger datasets.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BEATRICE: Bayesian Fine-mapping from Summary Datausing Deep Variational Inference 96%
- Gemini: Memory-efficient integration of hundreds of gene networks with high-order pooling 96%
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 96%
Similar papers in this journal
- Learning Genetic Perturbation Effects with Variational Causal Inference 97%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 97%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.