Back

Using machine learning to predict organismal growth temperatures from protein primary sequences

Sauer, D.; Wang, D.-N.

2019-06-21 bioinformatics
10.1101/677328 bioRxiv
Show abstract

BackgroundThe link between protein or nucleic acid sequence and biochemical or organismal phenotype is essential for understanding the molecular mechanisms of evolution, reverse ecology, and designing proteins and genes with specific properties. However, it is difficult to practically make use of the relationship between sequence and phenotype due to the complex relationship between sequence and folding or activity. ResultsHere, we predict the originating species optimal growth temperatures of individual protein sequences using trained machine learning models. Both multilayer perceptron and k Nearest Neighbor regression outperformed linear regression could predict the originating species optimal growth temperature from protein sequences, achieving a root mean squared error of 3.6 {degrees}C. Similar machine learning models could predict organismal optimal growth pH and oxygen tolerance, and the quantitative properties of individual proteins or nucleic acids. ConclusionsUsing multilayer perceptron and k Nearest Neighbor regressions, we were able to build models specific to individual protein or nucleic acid families that can predict a variety of quantitative phenotypes. This methodology will be useful the in silico screening of individual mutations for particular properties, and also effective in the predicting the phenotypes of uncharacterized biological sequences and organisms.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.