FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling
Zhu, J.; Shi, Y.; Bi, R.; Jin, P.; Liu, C.; Zhang, Z.; Huang, H.; Guo, Z.; Hu, P.; Ju, F.; Huang, L.; Tai, X.; Li, C.; Gao, K.; Wei, X.; Xia, H.; Zhang, J.; Min, Y.; Wang, Z.; Wang, Y.; He, L.; Liu, H.; Qin, T.
Show abstract
Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-only language models (e.g., ESM-2) capture sequence semantics at scale but lack structural grounding. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting evolutionary couplings, but their reliance on homologous sequences makes them less reliable in highly mutated or alignment-sparse regimes. We present FlexRibbon, a pretrained protein model that jointly learns from amino acid sequences and three-dimensional structures. Our pretraining strategy combines masked language modeling with diffusion-based denoising, enabling bidirectional sequence-structure learning without requiring MSAs. Trained on both experimentally resolved structures and AlphaFold 2 predictions, FlexRibbon captures global folds as well as flexible conformations critical for biological function. Evaluated across diverse tasks spanning interface design, intermolecular interaction prediction, and protein function prediction, FlexRibbon establishes new state-of-the-art performance on 12 different tasks, with particularly strong gains in mutation-rich settings where MSA-based methods often struggle.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.