CLIMBS: assessing Carbohydrate-Protein interactions through a graph neural network classifier using synthetic negative data
Luo, Y.; Parmeggiani, F.
Show abstract
I.Carbohydrate-protein interactions are essential for biological processes, such as cellular signaling and metabolism, and represent a large pool of untapped targets for diagnostics and therapeutics. However, current design and prediction methods fail to accurately evaluate the affinity and specificity of proteins for carbohydrates such as glucose and galactose. Here, we describe a machine learning classifier, named CLIMBS, as a novel scoring method for protein-carbohydrate interactions and train it on crystal structures and synthetic data from unsuccessfully designed binders, to effectively assess if carbohydrate-protein complexes represent realistic, native like structures. Compared to other methods, CLIMBS has outstanding accuracy, excellent carbohydrate specificity, sub-second runtime per sample, minimal bias towards either negative or positive samples, and can be employed to improve selection of successful docking and design models of carbohydrate-protein complexes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PROTACable is an Integrative Computational Pipeline of 3-D Modeling and Deep Learning to Automate the De Novo Design of PROTACs 96%
- CENsible: Interpretable Insights into Small-Molecule Binding with Context Explanation Networks 95%
- To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2 95%
- OPUS-Rota4: A Gradient-Based Protein Side-Chain Modeling Framework Assisted by Deep Learning-Based Predictors 94%
- MutateX: an automated pipeline for in-silico saturation mutagenesis of protein structures and structural ensembles 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.