FEDKEA: Enzyme function prediction with a large pretrained protein language model and distance-weighted k-nearest neighbor
Zheng, L.; Li, B.; Xu, S.; Chen, J.; Liang, G.
Show abstract
Recent advancements in sequencing technologies have led to the identification of a vast number of hypothetical proteins, surpassing current experimental capabilities for annotation. Enzymes, crucial for diverse biological functions, have garnered significant attention; however, accurately predicting enzyme EC numbers for proteins with unknown functions remains challenging. Here, we introduce FEDKEA, a novel computational method that integrates ESM-2 and distance-weighted KNN (k-nearest neighbor) to enhance enzyme function annotation. FEDKEA first employs a fine-tuned ESM-2 model with four fully connected layers to distinguish from other proteins. For predicting EC numbers, it adopts a hierarchical approach, utilizing distinct models and training strategies across the four EC number levels. Specifically, the classification of the first EC number level utilizes a fine-tuned ESM-2 model with three fully connected layers, while transfer learning with embeddings from this model supports the second and third-level tasks. The fourth-level classification employs a distance-weighted KNN model. Compared to existing tools such as CLEAN and ECRECer, two state-of-the-art computational methods, FEDKEA demonstrates superior performance. We anticipate that FEDKEA will significantly advance the prediction of enzyme functions for uncharacterized proteins, thereby impacting fields such as genomics, physiology and medicine. FEDKEA is easy to install and currently available at: https://github.com/Stevenleizheng/FEDKEA
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 97%
- GOBoost: Leveraging Long-Tail Gene Ontology Terms for Accurate Protein Function Prediction 97%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 97%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
Similar papers in this journal
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 95%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 96%
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 94%
- DeepHE: Accurately Predicting Human Essential Genes based on Deep Learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.