Back

Boltz-Perturb: Improving Diversity and Accuracy in Protein-Ligand Co-Folding through Training-Free Conditioning Perturbation

Jung, H.; Lee, B.; Cheng, A. C.

2026-08-05 biophysics
10.64898/2026.08.05.742877 bioRxiv
Show abstract

Protein-ligand co-folding models hold promise in structure-based drug discovery and small molecule interaction prediction, but often fail in predicting correct small molecule binding poses. We present Boltz-Perturb, a framework for addressing this through perturbing model conditioning signals during model inference, and show that such perturbations improve correct ligand binding mode predictions. We first show with true-coordinate injection experiments that the models learned energy landscape contains correct binding-mode basins, allowing us to reframe the problem as one of sampling deficiency. We then introduce two inference-time perturbation strategies, Token Bias Perturbation (TBP) and Token Conditioning Perturbation (TCP), which increase exploration of alternative binding poses. Across diverse protein-ligand systems, TCP improves top-20 oracle success rates by 2.6 to 7.8 fold. Boltz-Perturb attains higher oracle success rates compared to the Boltz-2 high diffusion temperature variant while requiring over 75% less compute. To our knowledge, this is the first systematic perturbation analysis of a co-folding architecture for small-molecule binding mode diversity. We demonstrate that inference-time perturbations can unlock latent structural diversity in generative co-folding models and improve protein-ligand predictions without costly retraining.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.