EvoPool: Evolution-Guided Pooling of Protein Language Model Embeddings
NaderiAlizadeh, N.; Singh, R.
Show abstract
Protein language models (PLMs) encode amino acid sequences into residue-level embeddings that must be pooled into fixed-size representations for downstream protein-level prediction tasks. Although these embeddings implicitly reflect evolutionary constraints, existing pooling strategies operate on single sequences and do not explicitly leverage information from homologous sequences or multiple sequence alignments. We introduce EvoPool, a self-supervised pooling framework that integrates evolutionary information from homologs directly into aggregated PLM representations using optimal transport. Our method constructs a fixed-size evolutionary anchor from an arbitrary number of homologous sequences and uses sliced Wasserstein distances to derive query protein embeddings that are geometrically informed by homologous sequence embeddings. Experiments across multiple state-of-the-art PLM families on the ProteinGym benchmark show that EvoPool consistently outperforms standard pooling baselines for variant effect prediction, demonstrating that explicit evolutionary guidance substantially enhances the functional utility of PLM representations. Our implementation code is available at https://github.com/navid-naderi/EvoPool.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 97%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 95%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.