Structure-informed Language Models Are Protein Designers
Zheng, Z.; Deng, Y.; Xue, D.; Zhou, Y.; YE, F.; Gu, Q.
Show abstract
This paper demonstrates that language models are strong structure-based protein designers. We present LM-DO_SCPLOWESIGNC_SCPLOW, a generic approach to reprogramming sequence-based protein language models (pLMs), that have learned massive sequential evolutionary knowledge from the universe of natural protein sequences, to acquire an immediate capability to design preferable protein sequences for given folds. We conduct a structural surgery on pLMs, where a lightweight structural adapter is implanted into pLMs and endows it with structural awareness. During inference, iterative refinement is performed to effectively optimize the generated protein sequences. Experiments show that LM-DO_SCPLOWESIGNC_SCPLOW improves the state-of-the-art results by a large margin, leading to 4% to 12% accuracy gains in sequence recovery (e.g., 55.65%/56.63% on CATH 4.2/4.3 single-chain benchmarks, and >60% when designing protein complexes). We provide extensive and in-depth analyses, which verify that LM-DO_SCPLOWESIGNC_SCPLOW can (1) indeed leverage both structural and sequential knowledge to accurately handle structurally non-deterministic regions, (2) benefit from scaling data and model size, and (3) generalize to other proteins (e.g., antibodies and de novo proteins).
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cross-Modality and Self-Supervised Protein Embedding for Compound-Protein Affinity and Contact Prediction 97%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 96%
- ProtMamba: a homology-aware but alignment-free protein state space model 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- All-Atom Protein Sequence Design using Discrete Diffusion Models 97%
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 95%
- Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.