Back

Rapid sequence-based screening of structure-disrupting protein mutations

Oh, J.; Qian, X.; Yoon, B.-J.

2026-02-25 bioinformatics
10.64898/2026.02.24.707693 bioRxiv
Show abstract

Recent advances in AI-based protein structure prediction have dramatically reduced the cost of obtaining three-dimensional protein models and have become integral to modern protein engineering workflows. However, full structure prediction remains computationally prohibitive in high-throughput settings for mutation-based protein engineering, where thousands of candidate variants may need to be evaluated. In many such scenarios, the primary objective is not to resolve the complete structure of a candidate mutant, but rather to identify whether the introduced mutations are likely to induce substantial structural changes for rapid down-selection of candidates that conserve the wildtype structure. Protein language models (PLMs), trained solely on unlabeled natural protein sequence data, are known to encode rich structural information within their hidden representations. Motivated by this observation, we investigate a range of sequence-based ranking metrics derived from PLMs as efficient surrogates for structural deformation prediction. Through systematic evaluation across multiple proteins, mutation regimes, and structure-prediction backbones, we show that embedding distance--particularly the L1 distance between representations--provides a robust and computationally efficient signal for identifying structure-disrupting mutations. Our results demonstrate that sequence-level screening can substantially reduce the need for expensive structure prediction while preserving sensitivity to large structural perturbations, thereby providing the means to significantly speed up mutation-based protein design.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.