Controllable Protein Design by Prefix-Tuning Protein Language Models
Luo, J.; Liu, X.; Li, J.; Chen, Q.; Chen, J.
Show abstract
The design of novel proteins with tailored functionalities, particularly in drug discovery and vaccine development, presents a transformative approach to addressing pressing biomedical challenges. Inspired by the remarkable success of pre-trained language models in natural language processing (NLP), protein language models (ProtLMs) have emerged as powerful tools in advancing protein science. While NLP leverages flexible text-based control tags to prompt language model generation, the restricted amino acid space (limited to 20 residues) imposes inherent constraints on achieving analogous controllability. In this study, we propose PrefixProt, a framework for controllable protein design that employs prefix tuning to learn virtual tokens as control tags. These virtual tokens are adaptively tailored to diverse protein properties through a data-driven manner and can be combinatorially integrated to enable multi-objective control over protein generation. The effectiveness of PrefixProt was validated through extensive experiments encompassing both protein structure design (e.g. alpha-helix or beta-sheet topologies) and protein function design (e.g. antimicrobial or anticancer peptide activities). Benchmark results demonstrate that prefix virtual tokens efficiently guide the pre-trained ProtLM by optimizing a smaller number of trainable parameters, out-performing other parameter-efficient fine-tuning methods and text-guided ProtLMs, particularly in scenarios with limited data availability. More importantly, the compositional flexibility of virtual tokens facilitates the generation of proteins with multiple target properties, substantially expanding the scope of design possibilities. By harmonizing controllability, efficiency and generalizability, PrefixProt establishes a robust framework for de novo protein design, with promising applications in drug discovery and biomedicine. Availability and implementationThe models and code are available at: https://github.com/chen-bioinfo/PrefixProt
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- InversePep: Diffusion-Driven Structure-Based Inverse Folding for Functional Peptides 96%
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 96%
- Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy 96%
Similar papers in this journal
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 97%
- PepFlow: direct conformational sampling from peptide energy landscapes through hypernetwork-conditioned diffusion 96%
- Generalized Biological Foundation Model with Unified Nucleic Acid and Protein Language 94%
Similar papers in this journal
Similar papers in this journal
- A Multi-Property Optimizing Generative Adversarial Network for de novo Antimicrobial Peptide Design 97%
- Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework 96%
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.