Back

Evolutionary profile enhancement improves protein function annotation for remote homologs

Dai, S.; Luo, J.; Luo, Y.

2026-03-04 bioinformatics
10.64898/2026.03.03.709280 bioRxiv
Show abstract

Accurate annotation of protein function is essential for understanding biological processes, yet this remains challenging for proteins lacking characterized homologs or belonging to underrepresented functional classes. Although machine learning approaches have become the gold standard for automated function prediction, they often perform poorly on out-of-distribution samples with low sequence identity to training proteins with known annotations. We propose EPERep, an evolutionary input enhancement strategy that leverages the vast space of unannotated protein sequences to improve the prediction of the functions of underrepresented proteins. Our key insight is that, even if a query protein has insufficient similarity to annotated proteins for direct annotation transfer, a wider range of similar unannotated sequences can be identified to facilitate better representation learning. Inspired by profile-based sequence search methods, EPERep incorporates homologous sequences as contextual input to refine the representations of individual proteins from pre-trained protein language models, effectively constructing a pLM-based profile for each query protein. Across four major annotation benchmarks on EC numbers, structural domains, Pfam families, and Gene Ontology predictions, EPERep consistently outperforms strong ML and sequence-alignment baselines. Gains are most pronounced for proteins from rare functional classes, with few or no labeled homologs, and for sequences exhibiting remote homology to the training distribution. These results demonstrate that evolutionary input enhancement provides a principled and scalable strategy for improving protein function prediction, particularly in long-tail and low-identity regimes.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.