Back

Constrained Generative Design Frameworks For Computational Discovery of Target-Specific DARPin Candidates

Pourbaghi, M.; Elemento, O.; Bradbury, M. S.

2026-08-13 bioinformatics
10.64898/2026.08.07.743551 bioRxiv
Show abstract

Applying unconstrained generative protein models to fixed structural scaffolds can produce systematic design artifacts, including a "Glycine Trap" characterized by the enrichment of glycine at structurally incompatible positions. Furthermore, optimizing sequences against artificial rigid-body docking geometries induces reward-hacking and severe geometric hallucinations. In addition, the highly conserved designed ankyrin repeat protein, or DARPin, scaffold can obscure defects at the engineered binding interface, causing AlphaFold2-Multimer (AF2) to predict nonfunctional protein-target interactions with high confidence. To overcome these limitations, we developed DARPinMPNN, a scaffold-constrained computational pipeline for DARPin candidate discovery. Restricting sequence generation to a validated DARPin design space eliminated these failure modes. A state-aware chimeric multiple sequence alignment strategy was engineered and enabled AlphaFold2-Multimer (AF2) to serve as a high-throughput structural sieve, while AlphaFold 3 (AF3) provided independent structural validation of candidate binders. Using this framework, we identified mesothelin-targeting DARPin candidates with predicted structural confidences (champion ipTM = 0.83) approaching those of a structurally validated picomolar-affinity binder (G3 control, ipTM = 0.89). By revealing extensive discordance between AF2 and AF3 predictions, this work establishes a robust framework for identifying and prioritizing high-confidence DARPin candidates for experimental validation.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.