A Benchmarking Platform for Assessing Protein Language Models on Function-related Prediction Tasks
Cevrim, E.; Yigit, M. G.; Ulusoy, E.; Yilmaz, A.; Dogan, T.
Show abstract
Proteins play a crucial role in almost all biological processes, serving as the building blocks of life and mediating various cellular functions, from enzymatic reactions to immune responses. Accurate annotation of protein functions is essential for advancing our understanding of biological systems and developing innovative biotechnological applications and therapeutic strategies. To predict protein function, researchers primarily rely on classical homology-based methods, which use evolutionary relationships, and increasingly on machine learning (ML) approaches. Lately, protein language models (PLMs) have gained prominence; these models leverage specialised deep learning architectures to effectively capture intricate relationships between sequence, structure, and function. We recently conducted a comprehensive benchmarking study to evaluate diverse protein representations (i.e., classical approaches and PLMs) and discuss their trade-offs. The current work introduces the Protein Representation Benchmark - PROBE tool, a benchmarking framework designed to evaluate protein representations on function-related prediction tasks. Here, we provide a detailed protocol for running the framework via the GitHub repository and accessing our newly developed user-friendly web service. PROBE encompasses four core tasks: semantic similarity inference, ontology-based function prediction, drug target family classification, and protein-protein binding affinity estimation. We demonstrate PROBEs usage through a new use case evaluating ESM2 and three recent multimodal PLMs--ESM3, ProstT5, and SaProt--highlighting their ability to integrate diverse data types, including sequence and structural information. This study underscores the potential of protein language models in advancing protein function prediction and serves as a valuable tool for both PLM developers and users.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Attention-based approach to predict drug-target interactions across seven target superfamilies 98%
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 97%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 96%
Similar papers in this journal
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 96%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 96%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 97%
- From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction 95%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.