Back

GraphESMStable: A Deep Learning Framework for Protein Stability Prediction Fusing Pre-trained Sequence Models and Graph Neural Networks

Zhang, H.; Li, Z.; He, J.

2026-01-09 bioengineering
10.64898/2026.01.08.698524 bioRxiv
Show abstract

Predicting protein stability is fundamental, yet existing computational methods often face limitations in generalization, reliance on single data modalities, and challenges with complex multi-point mutations. To address these, we propose GraphESMStable, a novel deep learning framework for predicting protein mutation-induced thermal stability changes. GraphESMStable integrates rich evolutionary context from a frozen pre-trained protein sequence language model with fine-grained three-dimensional structural geometry captured by a trainable Graph Neural Network. A sophisticated residue-level cross-attention mechanism facilitates the deep fusion of these distinct modal representations. The framework features a dedicated prediction head capable of predicting an entire single-point mutation landscape in a single forward pass, and an Epistasis Decoder explicitly modeling non-additive effects for multi-point mutations. Trained exclusively on a large-scale dataset, GraphESMStable achieves state-of-the-art performance, outperforming baselines across a diverse suite of independent evaluation benchmarks. This includes superior generalization on multiple stability datasets, cross-metric generalization to thermal melting prediction, and a significant lead in predicting double mutation epistatic effects. Furthermore, it demonstrates robust performance in predicting human pathogenic mutation stability and achieves substantial improvements in low-sample fitness prediction tasks. Our ablation studies confirm the synergistic benefits of this cross-modal fusion. GraphESMStable represents a significant advancement towards building highly generalizable and efficient foundational models for protein stability prediction, offering broad applicability in protein design and biomedical research.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.