Design and experimental characterization of specificity-switching mutational paths of WW domains
Rehan, A.; Mauri, E.; Fernandez-De-Cossio-Diaz, J.; Brun, P.-G.; Monasson, R.; Ribezzi-Crivellari, M.; Cocco, S.
Show abstract
Specific interactions between proteins and other biomolecules are ubiquitous in cellular processes. How specificity is encoded in the protein sequence and can be modified through a minimal set of concerted mutations is a complex issue. In this work, we focus on the WW protein domain, whose variants specifically bind to different classes of proline-rich peptides. Combining unsupervised learning of homologous WW sequence data with Restricted Boltzmann Machines (RBM) and path-sampling methods, we design mutational paths of putative WW domains interpolating between two natural WW domains with either distinct or similar specificities. Sequences along the designed paths are then experimentally validated with high-throughput in-vitro binding assays against 3 peptides of different classes. The vast majority (93%) of intermediate sequences along the designed paths are responsive to the initial or/and final peptides. On the contrary, domains along scrambled paths, in which the same mutations are introduced in random order are not functional, emphasizing how successful design crucially depends on the ability to model epistatic interactions. Interestingly, switch in specificity between classes I and IV whose representative peptides bind to different pockets on the WW domain appears to be smooth, with intermediates displaying some level of binding cross-reactivity with all tested peptides. We finally show that the RBM paths share a high identity with internal nodes obtained from ancestral sequence reconstruction based on the seed WW domains. Significance StatementGenerative machine-learning models are nowadays used to design new protein sequences with desired functions. Here, we address a more demanding task: designing a full mutational path connecting two natural proteins with different binding specificities. We illustrate this problem with WW domains, a small protein unit capable of recognizing distinct classes of proline-rich peptides. We experimentally verify that most of the intermediate sequences along the designed path are functional and respond to the initial or/and final peptides. The designed sequences share significant homology with the sequences obtained as internal nodes of phylogenetic trees through ancestral sequence reconstruction.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From complete cross-docking to partners identification and binding sites predictions 96%
- Generative and interpretable machine learning for aptamer design and analysis of in vitro sequence selection 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.