Toward De Novo Protein Design from Natural Language
Dai, F.; Fan, Y.; Su, J.; Wang, C.; Han, C.; Zhou, X.; Liu, J.; Qian, H.; Wang, S.; Zeng, A.; Wang, Y.; Yuan, F.
10.1101/2024.08.01.606258 bioRxivShow abstract
Programming biological function--designing bespoke proteins that execute specific tasks on demand--is a foundational goal of molecular engineering. Yet, current protein design paradigms remain fundamentally limited, typically requiring either an existing protein to evolve from, or deep, family-specific expertise to guide the design process. Here we introduce Pinal, a generative model that overcomes this barrier by translating functional descriptions in natural language directly into diverse and active proteins. This capability is built upon a 16-billion-parameter foundation model trained on an unprecedented synthetic corpus of 1.7 billion protein-text pairs, enabling it to ground functional language in the biophysical principles of protein structure. To provide definitive experimental validation, we tasked Pinal with designing four proteins from distinct functional classes: a fluorescent protein, a polyethylene terephthalate hydrolase, an alcohol dehydrogenase, and a metabolic H-protein. Remarkably, all four designs were functionally active and the two Pinal-designed enzymes achieved catalytic turnover for their respective reactions. Notably, the Pinal-designed H-protein even surpassed its natural counterpart, exhibiting 1.7-fold higher performance. Our results establish that natural language can serve as a programmable instruction set for biology, democratizing protein design and shifting the paradigm from the incremental modification of existing molecules to the direct creation of function from a conceptual description.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProtMamba: a homology-aware but alignment-free protein state space model 98%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 97%
- Cross-Modality and Self-Supervised Protein Embedding for Compound-Protein Affinity and Contact Prediction 97%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.