A combinatorial mutational map of active non-native protein kinases by deep learning guided sequence design
Seki, K.; Guo, A. B.; Akpinaroglu, D.; Kortemme, T.
Show abstract
Mapping protein sequence-function landscapes has either been limited to small steps (only few mutations) or to sequences similar to those already explored by evolution to maintain activity. Here, we overcome both limitations by applying deep-learning guided redesign to a natural protein tyrosine kinase to generate novel, functional sequences with highly combinatorial mutations. Using cell-free assays, we measure the activities and concentrations of 537 redesigned sequences, which differ from the wild-type by an average of 37 mutations while retaining activity in 85% of variants. These sequences sample 436 unique mutations at 76 different positions throughout the kinase domain. A simple regression model identifies key sequence determinants of function and predicts the function of unseen sequences. Our approach demonstrates how integrating deep-learning guided redesign, functional measurement at scale, and interpretable computational modelling enables functional exploration of highly combinatorial and sparse sequence-function landscapes at mutational scales not possible before.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The pathogenic T42A mutation in SHP2 rewires the interaction specificity of its N-terminal regulatory domain 97%
- Parametrically guided design of beta barrels and transmembrane nanopores using deep learning 95%
- Crystal structure of a guanine nucleotide exchange factor encoded by the scrub typhus pathogen Orientia tsutsugamushi 95%
Similar papers in this journal
Similar papers in this journal
- Genetic encoding of 3-nitro-tyrosine reveals the impacts of 14-3-3 nitration on client binding and dephosphorylation 95%
- Learning Peptide Recognition Rules for a Low-Specificity Protein 95%
- Neutralizing antibodies targeting the SARS-CoV-2 receptor binding domain isolated from a naïve human antibody library 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.