Back

To Predict is to Design: Unlocking Generative Capabilities in All-Atom Structure Predictors via Geometric Score Distillation

Li, Y.; Su, Y.; Jing, Y.; Liu, T.

2025-12-26 biophysics
10.64898/2025.12.23.696242 bioRxiv
Show abstract

Current protein binder design largely relies on a decoupled paradigm: generating backbones via unconditioned diffusion followed by sequence filling or refilling with inverse folding models. This separation prevents the design process from accessing the holistic validation metrics of structure predictors during generation, wasting rich physical priors. While recent works like BindCraft have successfully inverted AlphaFold2 for protein design1, extending this inversion to state-of-the-art all-atom diffusion predictors (e.g., AlphaFold3, Boltz-2) remains a formidable challenge, particularly for modalities requiring non-standard residues such as cyclic peptides. In this work, we present DREAM (Differentiable Refinement via Energy-Anchored Manifolds), a model-agnostic framework that turns the passive predictive trajectory of diffusion models into an active, lucid design process. DREAM repurposes Boltz-22--a leading open-source all-atom predictor--via Geometric Score Distillation (GSD), a technique enabling explicit gradient-based optimization directly through the frozen diffusion network. Unlike previous methods constrained by standard amino acids, DREAM directly unlocks the models latent chemical vocabulary, allowing gradients to autonomously select the optimal building blocks up to 55 residue types (including D-amino acids and post-translational modifications) to minimize energy. We demonstrate this capability by designing cyclic peptide binders for diverse targets, including PD-L1, B7-H3, and the human -Opioid Receptor (hMOR). Our results suggest that the programmable design of chemically complex modalities is not a distant goal, but a latent capability of current all-atom predictors, waiting to be inverted. Ultimately, to predict is to design.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.