Refining Embedding-Based Binding Predictions by Leveraging AlphaFold2 Structures
Endres, L.; Olenyi, T.; Erckert, K. Y.; Weissenow, K.; Rost, B.; Littmann, M.
Show abstract
BackgroundIdentifying residues in a protein involved in ligand binding is important for understanding its function. bindEmbed21DL is a Machine Learning method which predicts protein-ligand binding on a per-residue level using embeddings derived from the protein Language Model (pLM) ProtT5. This method relies solely on sequences, making it easily applicable to all proteins. However, highly reliable protein structures are now accessible through the AlphaFold Protein Structure Database or can be predicted using AlphaFold2 and ColabFold, allowing the incorporation of structural information into such sequence-based predictors. ResultsHere, we propose bindAdjust which leverages predicted distance maps to adjust the binding probabilities of bindEmbed21DL to subsequently boost performance. bindAdjust raises the recall of bindEmbed21DL from 47{+/-}2% to 53{+/-}2% at a precision of 50% for small molecule binding. For binding to metal ions and nucleic acids, bindAdjust serves as a filter to identify good predictions focusing on the binding site rather than isolated residues. Further investigation of two examples shows that bindAdjust is in fact able to add binding predictions which are not close in sequence but close in structure, extending the binding residue predictions of bindEmbed21DL to larger binding stretches or binding sites. ConclusionDue to its simplicity and speed, the algorithm of bindAdjust can easily refine binding predictions also from other tools than bindEmbed21DL and, in fact, could be applied to any protein prediction task.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From complete cross-docking to partners identification and binding sites predictions 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 96%
- Protein Stability Prediction by Fine-tuning a Protein Language Model on a Mega-scale Dataset 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.