NanoBERTa-ASP: Predicting Nanobody Binding Epitopes Based on a Pretrained RoBERTa Model
Li, S.; Meng, X.; Li, R.; Huang, B.; Wang, X.
Show abstract
Nanobodies, also known as VHH or single-domain antibodies, are a unique class of antibodies that consist only of heavy chains and lack light chains. Nanobodies possess the advantages of both small molecule drugs and conventional antibodies, making them a promising class of therapeutic biopharmaceuticals. Studying the characteristics of nanobody sequences can aid the development and design of nanobodies. An important challenge in this field is accurately predicting the binding sites between nanobodies and antigens. The binding site is the region where the nanobody interacts with the antigen. The precise prediction of these binding sites is crucial for comprehending the interaction mechanism between the nanobody and the antigen, facilitating the design of effective nanobodies, as well as gaining valuable insights into the properties of nanobodies. Nanobodies typically have smaller and more flexible binding sites than traditional antibodies, however predictive models trained on traditional antibodies may not be suitable for nanobodies. Moreover, the limited availability of antibodyderived nanobody datasets for deep learning poses challenges in constructing highly accurate predictive models that can generalize well to unseen data. To address these challenges, we propose a novel nanobody prediction model, named NanoBERTa-ASP (Antibody Specificity Prediction), which is specifically designed for predicting nanobody-antigen binding sites. The model adopts a training strategy more suitable for nanobodies by leveraging an advanced natural language processing (NLP) model called BERT (Bidirectional Encoder Representations from Transformers). The model also utilizes a masked language modeling approach to learn the contextual information of the nanobody sequence and predict its binding site. The results obtained from training NanoBERTa-ASP on nanobodies highlight its exceptional performance in Precision and AUC, underscoring its proficiency in capturing sequence information specific to nanobodies and accurately predicting their binding sites. Furthermore, our model can provide insights into the interaction mechanisms between nanobodies and antigens, contributing to a better understanding of nanobodies, as well as accelerating the development and design of nanobodies with potential therapeutic applications. To the best of our knowledge, NanoBERTa-ASP is the first nanobody language model that achieved high accuracy in predicting the binding sites. Github repositoryhttps://github.com/WangLabforComputationalBiology/NanoBERTa-ASP
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 96%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 96%
Similar papers in this journal
- Physical-aware model accuracy estimation for protein complex using deep learning method 96%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 95%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 95%
Similar papers in this journal
- nanoBERT: A deep learning model for gene agnostic navigation of the nanobody mutational space 96%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
- AbLang: An antibody language model for completing antibody sequences 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.