Back

HalluDesign: Protein Optimization and de novo Design via Iterative Structure Hallucination and Sequence Design

Fang, M.; Wang, C.; Shi, J.; Lian, F.; Jin, Q.; Wang, Z.; Zhang, Y.; Chen, P.; Cui, Z.; Wang, Y.; Zhang, Z.; Ke, Y.; Han, Q.; Cao, L.

2026-01-16 bioinformatics
10.1101/2025.11.08.686881 bioRxiv
Show abstract

Deep learning has revolutionized biomolecular modeling, enabling the prediction of diverse structures with atomic accuracy. However, leveraging the atomic-level precision of the structure prediction model for de novo design remains challenging. Here, we present HalluDesign, a general all-atom framework for protein optimization and de novo design, which iteratively update protein structure and sequence. HalluDesign harnesses the inherent hallucination capabilities of AlphaFold3-style structure prediction models and enables fine-tune free, forward-pass only sequence-structure co-optimization. Structure conditioning at different noise level in the structure prediction stage allows precise control over the sampling space, facilitating tasks from local and global protein optimization to de novo design. We demonstrate the versatility of this framework by optimizing suboptimal structures, rescuing previously unsuccessful designs, designing new biomolecular interactions and generating new protein structures from scratch. Experimental characterization of a diverse set of proteins--including protein binders for small molecules, a metal ion and proteins; antibody design of phosphorylation-specific peptide; and monomeric proteins--revealed high design success rates and excellent structural accuracy. Together, our comprehensive computational and experimental results highlight the broad utility of this framework. We anticipate that HalluDesign will further unlock the modeling and design potential of AlphaFold3-like models, enabling the robust creation of complex biomolecules for a wide range of biotechnological applications.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.