Back

Modeling the structure-conditioned sequence landscape for large-scale protein design with TriFlow

Srinivasan, H.; Yuan, R.; Cong, Q.; Zhou, J.

2025-12-02 bioinformatics
10.64898/2025.11.30.691458 bioRxiv
Show abstract

Generative models have revolutionized computational protein design, and the design of high-quality sequences given backbone structure is a critical component for success. Current state-of-the-art design pipelines utilize sequence design methods with local structural context and autoregressive generation. To improve efficiency and quality of sequence design, we developed TriFlow, a model that combines a RoseTTAFold-like three-track architecture for global structural context with discrete flow-matching for efficient few-step sequence generation. We trained TriFlow on a large dataset of interacting protein chains from PDB and interacting domains from AlphaFold protein structure Database to enrich its knowledge of natural protein/domain interfaces. While showing improvements across diverse benchmarks, TriFlows primary advance is in de novo binder design, where it boosts the in silico success rate of state-of-the-art design pipelines such as BindCraft. We demonstrated this by conducting a large-scale benchmark, generating and validating binders for over 500 diverse protein targets. By leveraging the model to explore the designed sequence landscape, we discovered that we can effectively highlight functional active sites, by contrasting constraints learned by the structure-conditioned model with natural evolutionary profiles. As a practical demonstration of its capabilities, we apply our pipeline to systematically design specific binders against human class I cytokines family, computationally optimizing for on-target affinity while minimizing off-target interactions, demonstrating that specificity also scales with inference time computational budget. TriFlow thus provides a robust framework for both large-scale protein engineering and for exploring the fundamental principles of the structure-conditioned sequence landscape.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.