Back

Effective sequence-to-expression prediction for membrane proteins using machine learning and computational protein design

Shen, Y.; Underhill, J.; Mulholland, A. J.; Oyarzun, D. A.; Curnow, P.

2025-09-28 bioengineering
10.1101/2025.09.25.678317 bioRxiv
Show abstract

The recombinant expression of integral membrane proteins is notoriously challenging. One way to address this challenge is via computational genotype-to-phenotype models that determine how particular sequence features correlate with protein expression levels. However, the potential of such approaches is yet to be fully realised, at least partly because so few expression datasets are available. Here, we study the sequence-to-expression relationships of a library of 12,248 variants of a specific membrane protein derived from combinatorial computational design. The major advantage of this approach lies in the controlled sequence diversity explored in design, making this new dataset directly compatible with lightweight off-the-shelf bioinformatic tools. The expression phenotype of the entire library is assessed in the widely-used recombinant host Escherichia coli. We employed a relatively small dataset of [~]2000 phenotyped sequences to train a sequence-to-expression predictor using supervised machine learning, which achieved high classification accuracy on held-out test sequences. This model was then used to infer the expression of >10,000 unmeasured sequences, and validation of the top predictions of both high and low expressers achieved 100% success rate. Using tools from explainable AI, we identified specific sequence positions and substitutions that are most important in dictating cellular expression levels. This analysis was validated by model-guided protein engineering that achieved an 8-fold increase in the purification yield of a poorly-expressing variant. Our results show that, at least for this controlled dataset, straightforward and interpretable machine learning can reveal the intrinsic sequence code for membrane protein expression.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.