Amalga: Designable Protein Backbone Generation with Folding and Inverse Folding Guidance
Chen, S.; Li, Z.; Zeng, X.; Ke, G.
Show abstract
Recent advances in deep learning enable new approaches to protein design through inverse folding and backbone generation. However, backbone generators may produce structures that inverse folding struggles to identify sequences for, indicating designability issues. We propose Amalga, an inference-time technique that enhances designability of backbone generators. Amalga leverages folding and inverse folding models to guide backbone generation towards more designable conformations by incorporating "folded-from-inverse-folded" (FIF) structures. To generate FIF structures, possible sequences are predicted from step-wise predictions in the reverse diffusion and further folded into new backbones. Being intrinsically designable, the FIF structures guide the generated backbones to a more designable distribution. Experiments on both de novo design and motif-scaffolding demonstrate improved designability and diversity with Amalga on RFdiffusion.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 97%
- To pack or not to pack: revisiting protein side-chain packing in the post-AlphaFold era 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 96%
Similar papers in this journal
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.