Back

Rescuing true protein binders from AI hallucinations via zero-shot, ensemble-driven statistical physics scoring

Chou, C.-H.; Hong, X.; Xu, J.

2026-05-18 bioinformatics
10.64898/2026.05.11.724213 bioRxiv
Show abstract

The advancement of deep generative models has facilitated de novo protein and antibody design, yet translation to experimental success is hindered by a high generation rate of structural decoys. Current affinity predictors and standard structural confidence metrics fail to reliably distinguish these AI hallucinations from true binders. Here, we present Sipobe-PPA, an affinity ranking framework that conceptualizes interacting protein interfaces as pseudo-ligands, evaluating them through an AI-driven statistical physics forcefield. Because this forcefield is trained exclusively on small-molecule interactions, Sipobe-PPA acts as a zero-shot physical evaluator for protein-protein interfaces, preventing the framework against the data leakage and memorization pitfalls that affect models trained directly on protein complex datasets. To capture the structural plasticity of binding interactions, Sipobe-PPA employs a conformational ensemble strategy, computing interaction scores across multiple AlphaFold3(AF3)-predicted structural states. Benchmarking on decoy-rich de novo datasets-including Bindcraft, Boltzgen, and the Germinal antibody dataset-demonstrates the significant improvement offered by this approach. In a real-world pipeline scenario simulating wet-lab constraints (pre-filtered by AF3 ipTM > 0.8 and pLDDT > 80), Sipobe-PPA achieved an 80% Hit Rate within its Top 5 predictions across the combined dataset, compared to 0% for physical baselines like Rosetta-{Delta}G. Notably, our structural ensemble averaging outperformed single-structure scoring, highlighting the necessity of modeling prediction diversity. By maximizing top-tier hit rates across diverse nanobody and de novo targets, Sipobe-PPA provides a scalable screening paradigm that bridges the gap between computational generation and wet-lab viability.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.