PhISCO: a simple method to infer phenotypes from protein sequences
Berthet, A. S. H.; Aptekmann, A. A.; bravo, j. t.; Sanchez, I. E.; Noguera, M. E.; Roman, E. A.
Show abstract
Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed PhISCO, Phenotype Inference from Sequence COmparisons, a simple algorithm that finds positions associated with any quantitative phenotype and predicts their values. From a few hundred sequences from four different protein families, we performed multiple sequence alignments and calculated per-position pairwise differences for both the sequence and the observed phenotypes. We found that from 3 to 10 positions, depending on the studied case, were enough to identify positions associated with the phenotypes and perform quantitative predictions of them. Here we show that these strong correlations can be found using individual positions while an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in observed phenotypes. We also show that although extremely simple, predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotype. The diversity of the explored systems makes PhISCO a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map 96%
- Disentangling the contribution of each descriptive characteristic of every single mutation to its functional effects 95%
- EVOLVE: A Web Platform for Evolutionary Phase Analysis and New Variant Exploration from Multi-Sequence Data 95%
Similar papers in this journal
- Improved prediction of stabilizing mutations in proteins by incorporation of mutational effects on ligand binding 96%
- Do Newly Born Orphan Proteins Resemble Never Born Proteins? A Study Using Three Deep Learning Algorithms 96%
- Prediction of protein assemblies by structure sampling followed by interface-focused scoring 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.