Improving functional protein generation via foundation model-derived latent space likelihood optimization
Guan, C.; Wan, F.; Torres, M. D. T.; de la Fuente-Nunez, C.
Show abstract
A variety of deep generative models have been adopted to perform de novo functional protein generation. Compared to 3D protein design, sequence-based generation methods, which aim to generate amino acid sequences with desired functions, remain a major approach for functional protein generation due to the abundance and quality of protein sequence data, as well as the relatively low modeling complexity for training. Although these models are typically trained to match protein sequences from the training data, exact matching of every amino acid is not always essential. Certain amino acid changes (e.g., mismatches, insertions, and deletions) may not necessarily lead to functional changes. This suggests that maximizing the training data likelihood beyond the amino acid sequence space could yield better generative models. Pre-trained protein large language models (PLMs) like ESM2 can encode protein sequences into a latent space, potentially serving as functional validators. We propose training functional protein sequence generative models by simultaneously optimizing the likelihood of training data in both the amino acid sequence space and the latent space derived from a PLM. This training scheme can also be viewed as a knowledge distillation approach that dynamically re-weights samples during training. We applied our method to train GPT- like models (i.e., autoregressive transformers) for antimicrobial peptide (AMP) and malate dehydrogenase (MDH) generation tasks. Computational experiments confirmed that our method outperformed various deep generative models (e.g., generative adversarial net, variational autoencoder, and GPT model without the proposed training strategy) on these tasks, demonstrating the effectiveness of our multi-likelihood optimization strategy.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 95%
Similar papers in this journal
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 96%
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 95%
- Evaluating generalizability of artificial intelligence models for molecular datasets 95%
Similar papers in this journal
- Designing diverse and high-performance proteins with a large language model in the loop 96%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 96%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 96%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.