Structured proteins are abundant in unevolved sequence space
Tretyachenko, V.; Vymetal, J.; Neuwirthova, T.; Vondrasek, J.; Fujishima, K.; Hlouchova, K.
Show abstract
Natural proteins represent numerous but tiny structure/function islands in a vast ocean of possible protein sequences, most of which has not been explored by either biological evolution or research. Recent studies have suggested this uncharted sequence space possesses surprisingly high structural propensity, but development of an understanding of this phenomenon has been awaiting a systematic high-throughput approach. Here, we designed, prepared, and characterized two combinatorial protein libraries consisting of randomized proteins, each 105 residues in length. The first library constructed proteins from the entire canonical alphabet of 20 amino acids. The second library used a subset of only 10 residues (A,S,D,G,L,I,P,T,E,V) that represent a consensus view of plausibly available amino acids through prebiotic chemistry. Our study shows that compact conformations resistant to proteolysis are (i) abundant (up to 40%) in random sequence space, (ii) independent of general Hsp70 chaperone system activity, and (iii) not granted solely by "late" and complex amino acid additions. The Hsp70 chaperone system effectively increases solubility and refoldability of the canonical alphabet but has only a minor impact on the "early" library. The early alphabet proteins are inherently more soluble and refoldable, possibly assisted by the cell-like environment in which these assays were performed. Our work indicates that both early and modern amino acids are predisposed to supporting protein structure (either in forms of oligomers or globular/molten globule structures) and that protein structure may not be a unique outcome of evolution.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Recombination of 2Fe-2S ferredoxins reveals differences in the inheritance of thermostability and midpoint potential 93%
- Cell-free synthesis of natural compounds from genomic DNA of biosynthetic gene clusters 93%
- LyGo: A platform for rapid screening of lytic polysaccharide monooxygenase production 93%
Similar papers in this journal
- Cracking Controls ATP Hydrolysis in the catalytic unit of a P-type ATPase 94%
- Triggering closure of a sialic acid TRAP transporter substrate binding protein through binding of natural or artificial substrates 93%
- What Strengthens Protein-Protein Interactions: Analysis and Applications of Residue Correlation Networks 93%
Similar papers in this journal
- Experimental characterization of in silico red-shift predicted iLOVL470T/Q489K and iLOVV392K/F410V/A426S mutants 93%
- Chemical targeting of the ATXN1 aa99-163 interaction site suppresses polyQ-expanded protein dimerization 92%
- Exploration of DPP-IV inhibitory peptide design rules assisted by deep learning pipeline that identifies restriction enzyme cutting site 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.