ProfileView: multiple probabilistic models resolve protein families functional diversity
Vicedomini, R.; Bouly, J.-P.; Laine, E.; Falciatore, A.; Carbone, A.
Show abstract
Sequence functional classification has become a critical bottleneck in understanding the myriad of protein sequences that accumulate in our databases. The great diversity of homologous sequences hides, in many cases, a variety of functional activities that cannot be anticipated. Their identification appears critical for a fundamental understanding of living organisms and for biotechnological applications. ProfileView is a sequence-based computational method, designed to functionally classify sets of homologous sequences. It relies on two main ideas: the use of multiple probabilistic models whose construction explores evolutionary information in available databases, and a new definition of a representation space where to look at sequences from the point of view of probabilistic models combined together. ProfileView classifies families of proteins for which functions should be discovered or characterised within known groups. We validate ProfileView on seven classes of widespread proteins, involved in the interaction with nucleic acids, amino acids and small molecules, and in a large variety of functions and enzymatic reactions. ProfileView agrees with the large set of functional data collected for these proteins from the literature regarding the organisation into functional subgroups and residues that characterize the functions. Furthermore, ProfileView resolves undefined functional classifications and extracts the molecular determinants underlying protein functional diversity, showing its potential to select sequences towards accurate experimental design and discovery of new biological functions. ProfileView proves to outperform three functional classification approaches, CUPP, PANTHER, and a recently developed neural network approach based on Restricted Boltzmann Machines. It overcomes time complexity limitations of the latter.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 95%
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 95%
- Decoding proteome functional information in model organisms using protein language models. 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.