BindCORE: Biophysical Ensemble Learning for Predicting Interaction Sites in Intrinsically Disordered Regions
Buton, N.; Piochi, L. F.; Khakzad, H.
Show abstract
Intrinsically disordered proteins and regions (IDPs/IDRs) mediate diverse cellular functions through binding segments whose functional properties are encoded in dynamic conformational ensembles rather than a single static state. Existing predictors of linear interacting peptides (LIPs) and molecular recognition features (MoRFs) rely primarily on sequence-derived features, leaving ensemble-level biophysical properties largely unexplored. Here, we introduce BindCORE, an ensemble-aware deep learning framework that integrates global, local, and pairwise biophysical descriptors to predict interaction sites within IDRs. These features are processed through a multi-scale architecture that enables information exchange between sequence- and ensemble-based global, local, and pairwise information. Across established LIP and MoRF benchmarks, BindCORE consistently improves performance over sequence-based baselines, demonstrating the predictive signals of ensemble-derived properties beyond sequence-based representations alone. Feature-attribution analyses reveal that pairwise descriptors are the dominant contributors to prediction, while solvent accessibility, backbone dihedral entropy, and global geometric properties provide complementary information. Feature-importance rankings vary substantially across ensemble flavours, indicating that different conformational generators encode distinct biophysical signatures of interaction-site propensity. Together, our results show that conformational ensembles contain interpretable determinants of LIP and MoRF binding residues and establish BindCORE as a general framework for incorporating biophysical information into the prediction of functional regions in intrinsically disordered proteins. BindCORE is freely available as a ready-to-use Google Colab notebook at https://gitlab.inria.fr/delta/bindcore.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BABAPPAlign: A Multiple Sequence Alignment Engine with a Learned Residue-Level Scoring Function 94%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 94%
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 94%
Similar papers in this journal
- Hybrid Deep Learning with Protein Language Models and Dual-Path Architecture for Predicting IDP Functions 97%
- InversePep: Diffusion-Driven Structure-Based Inverse Folding for Functional Peptides 95%
- ProDualNet: Dual-Target Protein Sequence Design Method Based on Protein Language Model and Structure Model 95%
Similar papers in this journal
- Efficient protein structure generation with sparse denoising models 94%
- Conditional Diffusion with Locality-Aware Modal Alignment for Generating Diverse Protein Conformational Ensembles 93%
- Predicting RNA 3D structure and conformers using a pre-trained secondary structure model and structure-aware attention 93%
Similar papers in this journal
- ConforFold Recovers Alternative Protein Conformations Beyond MSA Subsampling 96%
- Assessing the relation between protein phosphorylation, AlphaFold3 models and conformational variability 95%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.