EZpred: improving deep learning-based enzyme function prediction using unlabeled sequence homologs
Zhang, C.; Liu, Q.; Freddolino, L.
Show abstract
Features extracted from sequence homologs significantly enhance the accuracy of deep learning-based protein structure prediction. Indeed, models such as AlphaFold, which extracts features from sequence homologs, generally produce more accurate protein structures compared to single sequence-based methods like ESMfold. In contrast, features from sequence homologs are seldom employed for deep learning-based protein function prediction. Although a small number of models also incorporate function labels from sequence homologs, they cannot utilize features extracted from sequence homologs that lack function labels. To address this gap, we propose EZpred, which is the first deep learning model to use unlabeled sequence homologs for protein function prediction. Starting with the target sequence and homologs identified by MMseqs2, EZpred extracts sequence features using the ESMC protein language model. These features are then fed into a deep learning model to predict the Enzyme Commission (EC) numbers of the target protein. For 753 enzymes, the F1-score of EZpred EC number prediction is 4% higher than a similar model that does not use sequence homologs and at least 10% higher that state-of-the-art EC number prediction models. These results demonstrate the strong positive impact of sequence homologs in deep learning-based enzyme function prediction. Significance StatementMultiple sequence alignment (MSA) of homologous sequences is the most important source of input feature for deep learning-based protein structure prediction. Yet, this kind of feature is rarely used for protein function prediction. We propose the first deep learning model that significantly improves protein function prediction using features from sequence homologs without functional labels. This shows the utility of an important set of features that are often overlooked by previous studies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 96%
- TemStaPro: protein thermostability prediction using sequence representations from protein language models 95%
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 95%
Similar papers in this journal
- DeepSS2GO: protein function prediction from secondary structure 97%
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 96%
- Rossmann-toolbox: a deep learning-based protocol for the prediction and design of cofactor specificity in Rossmann-fold proteins 95%
Similar papers in this journal
- PremPS: Predicting the Effects of Single Mutations on Protein Stability 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 94%
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 94%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 94%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.