The genetic architecture of the human bZIP family
Bendel, A. M.; Eichenbergers, B.; Kempf, G.; Klein, D.; Shimada, K.; Franco, M.; Jaaks-Kraatz, S.; Durdu, S.; Herrmann, V.; Roth, G.; Soneson, C.; Smallwood, S.; Thomae, N. H.; Schubeler, D.; Stadler, M. B.; Zenke, F.; Diss, G.
Show abstract
Generative biology holds the promise to transform our ability to design and understand living systems by creating novel proteins, pathways, and organisms with tailored functions that address challenges in medicine, sustainability, and technology. However, training generative models requires large quantities of data that captures the genetic architecture of protein function in all its complexity, but these are currently scarce. Here, we systematically mutagenized all 54 human basic-leucine zipper (bZIP) domains and quantified their interactions with each other using bindingPCA, a quantitative deep mutational scanning assay. This resulted in [~]2 million interaction measurements, capturing the effect of all single amino acid substitutions at each of the 35 interfacial positions. We found that mutation effects are largely additive in the vicinity of each wild-type bZIPs, but diverge across the family, indicating strong context dependency. A global additive thermodynamic model provided moderate prediction of mutation effects, while individual models per bZIP achieved higher performance, supporting a model of local simplicity and global complexity. Our results therefore suggest that the genetic architecture of protein function is more complex than previously anticipated, which could hinder predictability. However, a convolutional neural network trained on this dataset could accurately predict binding scores from sequence alone. Furthermore, the model enabled the design of synthetic bZIPs with high target specificity, demonstrating practical applicability for bioengineering purposes. Our study shows that capturing family-wide diversity is essential to reveal context dependencies and train accurate quantitative models of protein-protein interactions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Epistasis facilitates functional evolution in an ancient transcription factor 97%
- Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator 96%
- Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins 95%
Similar papers in this journal
- Evolutionary-Scale Enzymology Enables Biochemical Constant Prediction Across a Multi-Peaked Catalytic Landscape 96%
- Short tandem repeats bind transcription factors to tune eukaryotic gene expression 95%
- Shifting mutational constraints in the SARS-CoV-2 receptor-binding domain during viral evolution 95%
Similar papers in this journal
Similar papers in this journal
- Active learning of enhancer and silencer regulatory grammar in a developing neural tissue 95%
- Rugged fitness landscapes minimize promiscuity in the evolution of transcriptional repressors 95%
- Synthetic coevolution reveals adaptive mutational trajectories of neutralizing antibodies and SARS-CoV-2 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.