Evolutionary profile enhancement improves protein function annotation for remote homologs
Dai, S.; Luo, J.; Luo, Y.
Show abstract
Accurate annotation of protein function is essential for understanding biological processes, yet this remains challenging for proteins lacking characterized homologs or belonging to underrepresented functional classes. Although machine learning approaches have become the gold standard for automated function prediction, they often perform poorly on out-of-distribution samples with low sequence identity to training proteins with known annotations. We propose EPERep, an evolutionary input enhancement strategy that leverages the vast space of unannotated protein sequences to improve the prediction of the functions of underrepresented proteins. Our key insight is that, even if a query protein has insufficient similarity to annotated proteins for direct annotation transfer, a wider range of similar unannotated sequences can be identified to facilitate better representation learning. Inspired by profile-based sequence search methods, EPERep incorporates homologous sequences as contextual input to refine the representations of individual proteins from pre-trained protein language models, effectively constructing a pLM-based profile for each query protein. Across four major annotation benchmarks on EC numbers, structural domains, Pfam families, and Gene Ontology predictions, EPERep consistently outperforms strong ML and sequence-alignment baselines. Gains are most pronounced for proteins from rare functional classes, with few or no labeled homologs, and for sequences exhibiting remote homology to the training distribution. These results demonstrate that evolutionary input enhancement provides a principled and scalable strategy for improving protein function prediction, particularly in long-tail and low-identity regimes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizable and scalable protein stability prediction with rewired protein generative models 98%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 97%
- Structure-Based Function Prediction using Graph Convolutional Networks 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.