Incorporating Pre-training Paradigm for Antibody Sequence-Structure Co-design
Gao, K.; Wu, L.; Zhu, J.; Peng, T.; Xia, Y.; He, L.; Xie, S.; Qin, T.; Liu, H.; He, K.; Liu, T.-Y.
Show abstract
Antibodies are versatile proteins that can bind to pathogens and provide effective protection for human body. Recently, deep learning-based computational antibody design has attracted popular attention since it automatically mines the antibody patterns from data that could be complementary to human experiences. However, the computational methods heavily rely on the high-quality antibody structure data, which is quite limited. Besides, the complementarity-determining region (CDR), which is the key component of an antibody that determines the specificity and binding affinity, is highly variable and hard to predict. Therefore, data limitation issue further raises the difficulty of CDR generation for antibodies. Fortunately, there exists a large amount of sequence data of antibodies that can help model the CDR and alleviate the reliance on structured data. By witnessing the success of pre-training models for protein modeling, in this paper, we develop an antibody pre-trained language model and incorporate it into the (antigen-specific) antibody design model in a systemic way. Specifically, we first pre-train an antibody language model based on the sequence data, then propose a one-shot way for sequence and structure generation of CDR to avoid the heavy cost and error propagation from an autoregressive manner, and finally leverage the pre-trained antibody model for the antigen-specific antibody generation model with some carefully designed modules. Through various experiments, we show that our method achieves superior performance over previous baselines on different tasks, such as sequence and structure generation, antigen-binding CDR-H3 design.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 96%
- Identifying B-cell epitopes using AlphaFold2 predicted structures and pretrained language model 96%
- Learning Context-aware Structural Representations to Predict Antigen and Antibody Binding Interfaces 95%
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.