Protein Sequence Design by Entropy-based Iterative Refinement
Zhou, X.; Chen, G.; Ye, J.; Wang, E.; Zhang, J.; Mao, C.; Li, Z.; Hao, J.; Huang, X.; Tang, J.; Heng, P. A.
Show abstract
Inverse Protein Folding (IPF) is an important task of protein design, which aims to design sequences compatible with a given backbone structure. Despite the prosperous development of algorithms for this task, existing methods tend to leverage limited and noisy residue environment when generating sequences. In this paper, we develop an iterative sequence refinement pipeline, which can refine the sequence generated by existing sequence design models. It selects and retains reliable predictions based on the models confidence in predicted distributions, and decodes the residue type based on a partially visible environment. The proposed scheme can consistently improve the performance of a number of IPF models on several sequence design benchmarks, and increase sequence recovery of the SOTA model by up to 10%. We finally show that the proposed model can be applied to redesign Transposon-associated transposase B. 8 variants exhibit improved gene editing activity among the 20 variants we proposed. Our code and a demo of the refinement pipeline are provided in the online colab.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 97%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 97%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 96%
Similar papers in this journal
- AnglesRefine: refinement of 3D protein structures using Transformer based on torsion angles 96%
- GAT-HiC: Efficient Reconstruction of 3D Chromosome Structure via Residual Graph Attention Neural Networks 96%
- Learning universal knowledge graph embedding for predicting biomedical pairwise interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.