All-Atom Protein Generation with Latent Diffusion
Lu, A. X.; Yan, W.; Robinson, S. A.; Kelow, S.; Yang, K. K.; Gligorijevic, V.; Cho, K.; Bonneau, R.; Abbeel, P.; Frey, N. C.
Show abstract
AbstractO_ST_ABSPurposeC_ST_ABSDesigning proteins with atomic-level functional control remains a central challenge in de novo design, exacerbated by limitations in structural data availability. MethodsWe introduce PLAID (Protein Latent Induced Diffusion), a generative model that efficiently co-generates discrete sequence and all-atom structure by sampling directly in the shared sequence-structure latent of a pretrained sequence-to-structure predictor. Unlike existing de novo generative models, PLAID trains its diffusion model exclusively on sequences, expanding effective training corpora by 2-4 orders of magnitude relative to structural databases. Using classifier-free guidance, PLAID supports controllable generation based on function (Gene Ontology) and organism keywords. ResultsIn silico, PLAID can unconditionally generate all-atom structures without explicit structural supervision during diffusion model training. Function conditioned proteins can recapitulates catalytic side-chain positions for residues at non-adjacent positions, and transmembrane proteins with expected hydrophobicity patterns and predicted topologies. We further experimentally validate that PLAID can be prompted to generate heme binding proteins with high sequence novelty. ConclusionOverall, PLAID unifies sequence-scale training with atomic-level generation, enabling more precise functional control in protein design. Code and weights: github.com/amyxlu/plaid.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- All-Atom Protein Sequence Design using Discrete Diffusion Models 96%
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 93%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.