Back

From Atoms to Fragments: A Coarse Representation for Functional and Efficient Protein Design

Castorina, L. V.; Wood, C. W.; Subr, K.

2025-03-19 bioinformatics
10.1101/2025.03.19.644162 bioRxiv
Show abstract

MotivationAlthough deep learning has accelerated protein design, current protein representations such as sequences or full-atom structures scale non-linearly with protein length. We propose a sparse and interpretable representation for proteins, based on evolutionarily conserved fragments. Specifically, we use a curated set of 40 functional and evolutionarily conserved fragments as an alphabet to build Fragment Graphs and Fragment Sets. These fragment-based representations are both lightweight and functionally informative, capturing up to 55% more variance using fewer than [Formula] of the dimensions required by traditional methods. ResultsOn a dataset of 215 functionally diverse proteins, our approach creates more coherent functional clusters than traditional sequence- and structure-based methods, even among proteins with [≤] 30% sequence identity. Fragment-based searches of protein databases achieve accuracies comparable to traditional methods, while using 90% fewer tokens per protein. These searches execute [~]68.7x faster than RMSD-based structural methods and [~]1.64x faster than sequence-based methods, even including fragment pre-processing overhead. Additionally, we show that our representation effectively guides RFDiffusion for protein backbone generation with functional recovery rates higher than 40%. In summary, our fragment-based representation offers a scalable and interpretable alternative for the next generation of protein design tools for backbone design, sequence design, and functional similarity searches within protein structure databases. Availabilityhttps://github.com/wells-wood-research/tessera (Documentation to be made available upon acceptance)

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.