Back

Structure-informed Language Models Are Protein Designers

Zheng, Z.; Deng, Y.; Xue, D.; Zhou, Y.; YE, F.; Gu, Q.

2023-02-03 bioinformatics
10.1101/2023.02.03.526917 bioRxiv
Show abstract

This paper demonstrates that language models are strong structure-based protein designers. We present LM-DO_SCPLOWESIGNC_SCPLOW, a generic approach to reprogramming sequence-based protein language models (pLMs), that have learned massive sequential evolutionary knowledge from the universe of natural protein sequences, to acquire an immediate capability to design preferable protein sequences for given folds. We conduct a structural surgery on pLMs, where a lightweight structural adapter is implanted into pLMs and endows it with structural awareness. During inference, iterative refinement is performed to effectively optimize the generated protein sequences. Experiments show that LM-DO_SCPLOWESIGNC_SCPLOW improves the state-of-the-art results by a large margin, leading to 4% to 12% accuracy gains in sequence recovery (e.g., 55.65%/56.63% on CATH 4.2/4.3 single-chain benchmarks, and >60% when designing protein complexes). We provide extensive and in-depth analyses, which verify that LM-DO_SCPLOWESIGNC_SCPLOW can (1) indeed leverage both structural and sequential knowledge to accurately handle structurally non-deterministic regions, (2) benefit from scaling data and model size, and (3) generalize to other proteins (e.g., antibodies and de novo proteins).

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.