Effective sequence-to-expression prediction for membrane proteins using machine learning and computational protein design
Shen, Y.; Underhill, J.; Mulholland, A. J.; Oyarzun, D. A.; Curnow, P.
Show abstract
The recombinant expression of integral membrane proteins is notoriously challenging. One way to address this challenge is via computational genotype-to-phenotype models that determine how particular sequence features correlate with protein expression levels. However, the potential of such approaches is yet to be fully realised, at least partly because so few expression datasets are available. Here, we study the sequence-to-expression relationships of a library of 12,248 variants of a specific membrane protein derived from combinatorial computational design. The major advantage of this approach lies in the controlled sequence diversity explored in design, making this new dataset directly compatible with lightweight off-the-shelf bioinformatic tools. The expression phenotype of the entire library is assessed in the widely-used recombinant host Escherichia coli. We employed a relatively small dataset of [~]2000 phenotyped sequences to train a sequence-to-expression predictor using supervised machine learning, which achieved high classification accuracy on held-out test sequences. This model was then used to infer the expression of >10,000 unmeasured sequences, and validation of the top predictions of both high and low expressers achieved 100% success rate. Using tools from explainable AI, we identified specific sequence positions and substitutions that are most important in dictating cellular expression levels. This analysis was validated by model-guided protein engineering that achieved an 8-fold increase in the purification yield of a poorly-expressing variant. Our results show that, at least for this controlled dataset, straightforward and interpretable machine learning can reveal the intrinsic sequence code for membrane protein expression.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pooled PPIseq: screening the SARS-CoV-2 and human interface with a scalable multiplexed protein-protein interaction assay platform 93%
- Improved yellow-green split fluorescent proteins for protein labeling and signal amplification 93%
- Promiscuous structural cross-compatibilities between major shell components of Klebsiella pneumoniae bacterial microcompartments 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.