Back

Accuracy and Scalability of Machine Learning Methods for Genotype-Phenotype Association Data

Collienne, K.; Zhang, L.; Gavryushkin, A.

2025-02-17 bioinformatics
10.1101/2025.02.13.638022 bioRxiv
Show abstract

Many machine learning methods can be applied to predicting phenotypes from genetic data. Which of these methods work best remains an open question, however. To answer this question, we propose to compare a variety of approaches ability to predict a simulated non-linear complex trait. Specifically, we evaluate these methods on their accuracy and scalability with respect to the amount training data available, the noise present in the data, the complexity of the simulated (trait) functions, and their ability to provide insight into the simulated trait. We then compare the best approach to state-of-the-art models in real data, predicting gout in the UK Biobank. We find that transformer encoders outperform all other methods in simulations, and perform comparably to the state-of-the-art with real data, with a promise to scale to significantly larger datasets.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.