Top-down design of protein nanomaterials with reinforcement learning
Lutz, I. D.; Wang, S.; Norn, C.; Borst, A. J.; Zhao, Y. T.; Dosey, A.; Cao, L.; Li, Z.; Baek, M.; King, N. P.; Ruohola-Baker, H.; Baker, D.
Show abstract
The multisubunit protein assemblies that play critical roles in biology are the result of evolutionary selection for function of the entire assembly, and hence the subunits in structures such as icosahedral viral capsids often fit together with remarkable shape complementarity1,2. In contrast, the large multisubunit assemblies that have been created by de novo protein design, notably the icosahedral nanocages used in a new generation of potent vaccines3-7, have been built by first designing symmetric oligomers with cyclic symmetry and then assembling these into nanocages while keeping the internal structure fixed8-14, which results in more porous structures with less extensive shape matching between the components. Such hierarchical "bottom-up" design approaches have the advantage that one interface can be designed and validated in the context of the cyclic oligomer building block15,16, but the disadvantage that the structural and functional features of the assemblies are limited by the properties of the predesigned building blocks. To overcome this limitation, we set out to develop a "top-down" reinforcement learning based approach to protein nanomaterial design in which both the structures of the subunits and the interactions between them are built up coordinately in the context of the entire assembly. We developed a Monte Carlo tree search (MCTS) method17,18 which assembles protein monomer structures in the context of an overall architecture guided by a loss function which enables specification of any desired overall structural properties such as shape and porosity. We demonstrate the power of the approach by designing hyperstable icosahedral assemblies more compact than any previously observed protein icosahedral structure (designed or naturally occurring), that have very low porosity and are robust to fusion and display of proteins as complex as influenza hemagglutinin. CryoEM structures of two designs are very close to the computational design models. Our top-down reinforcement learning approach should enable the design of a wide variety of complex protein nanomaterials by direct optimization of overall system properties.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 97%
- Accurate prediction of protein assembly structure by combining AlphaFold and symmetrical docking 97%
- Improved protein structure refinement guided by deep learning based accuracy estimation 97%
Similar papers in this journal
- qFit 3: Protein and ligand multiconformer modeling for X-ray crystallographic and single-particle cryo-EM density maps 96%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
Similar papers in this journal
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 97%
- TomoTwin: Generalized 3D Localization of Macromolecules in Cryo-electron Tomograms with Structural Data Mining 96%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
Similar papers in this journal
Similar papers in this journal
- A conserved glutathione binding site in poliovirus is a target for antivirals and vaccine stabilisation 96%
- Real-Time Structure Search and Structure Classification for AlphaFold Protein Models 95%
- Delineating organizational principles of the endogenous L-A virus by cryo-EM and computational analysis of native cell extracts 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.