BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings
Piao, H.; Boorla, V.; Santra, S.; Maranas, C. D.
Show abstract
MotivationReliable predictions of protein-protein binding affinities are essential for molecular biology and therapeutic discovery. However, most computational methods rely on three-dimensional structural models, which are often unavailable for many complexes. ResultsWe introduce BindPred, a structure-agnostic input framework that predicts affinities directly from amino acid sequences by combining embeddings from large protein language models with gradient boosting trees. On the PPB-Affinity benchmark, which comprises 11,919 diverse complexes, BindPred achieves a Pearson correlation coefficient of 0.86 in random split five-fold cross-validation, where the training and test sets share <30% global sequence identity. Ablation analysis indicates that evolutionary embeddings alone capture most of the predictive signals, while augmenting with physics-based energy terms from PyRosetta and BindCraft increases the correlation by only 0.01. A more stringent protein-level split that places entire protein families (wild-type and all mutants) exclusively in either training or testing sets, results in only a modest decline in performance, demonstrating robust generalization to novel interaction pairs. Because BindPred operates exclusively on sequence input, it enables rapid inference (approximately 3 million complexes per GPU (T4) hour), making proteome-scale screening computationally feasible. AvailabilityThe pretrained model and inference pipeline are available in a Google Colab notebook: BindPred Colab notebook. The training dataset, code, and model weights are available on the Hugging Face: BindPred Contactcostas@psu.edu Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
- MAHOMES II: A webserver for predicting if a metal binding site is enzymatic 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.