Back

FrustrAI-Seq: Scaling Local Energetic Frustration to the Protein Sequence Space

Leusch, J.-P.; Poley-Gil, M.; Fernandez-Martin, M.; Bordin, N.; Rost, B.; Parra, R. G.; Heinzinger, M.

2026-02-05 bioinformatics
10.64898/2026.02.03.703498 bioRxiv
Show abstract

Proteins fold into their native three-dimensional (3D) structures by navigating complex energy landscapes shaped by the biophysical and biochemical properties of their sequence. Once folded, some sequence positions (dubbed residues) remain locally frustrated, reflecting functional constraints incompatible with optimal packing. This local energetic frustration provides important insights into protein function and dynamics, but its analysis typically relies on structure-based energy calculations and remains energetically costly at scale. Here, we introduce an ultra-fast sequence-based prediction of local energetic frustration directly from protein sequences using embeddings from protein language models (pLMs). Our method, coined FrustrAI-Seq, enables proteome-wide frustration profiling in minutes ([~] 17 minutes for the entire human proteome on a single Nvidia H100 GPU) while retaining biologically relevant performance as shown for the -globin and {beta}-lactamase family. By eliminating the need for explicit structural or evolutionary information, this approach expands frustration analysis to protein regions and classes that were previously inaccessible, including intrinsically disordered regions and high-throughput de novo designed protein datasets. To support reproducibility and large-scale applications, we provide the largest freely available resource of precomputed local frustration scores to date ([~]106 proteins), along with model weights and complete training and inference code at: github.com/leuschjanphilipp/FrustrAI-Seq.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.