Establishing comprehensive quaternary structural proteomes from genome sequence
Catoiu, E. A.; Mih, N.; Lu, M.; Palsson, B. O.
Show abstract
A critical body of knowledge has developed through advances in protein microscopy, protein-fold modeling, structural biology software, availability of sequenced bacterial genomes, large-scale mutation databases, and genome-scale models. Based on these recent advances, we develop a computational framework that; i) identifies the oligomeric structural proteome encoded by an organisms genome from available structural resources; ii) maps multi-strain alleleomic variation, resulting in the structural proteome for a species; and iii) calculates the 3D orientation of proteins across subcellular compartments with residue-level precision. Using the platform, we; iv) compute the quaternary E. coli K-12 MG1655 structural proteome; v) use a dataset of 12,000 mutations to build Random Forest classifiers that can predict the severity of mutations; and, in combination with a genome-scale model that computes proteome allocation, vi) obtain the spatial allocation of the E. coli proteome. Thus, in conjunction with relevant datasets and increasingly accurate computational models, we can now annotate quaternary structural proteomes, at genome-scale, to obtain a molecular-level understanding of whole-cell functions. SignificanceAdvancements in experimental and computational methods have revealed the shapes of multi-subunit proteins. The absence of a unified platform that maps actionable datatypes onto these increasingly accurate structures creates a barrier to structural analyses, especially at the genome-scale. Here, we describe QSPACE, a computational annotation platform that evaluates existing resources to identify the best-available structure for each protein in a users query, maps the 3D location of actionable datatypes (e.g., active sites, published mutations) onto the selected structures, and uses third-party APIs to determine the subcellular compartment of all amino acids of a protein. As proof-of-concept, we deployed QSPACE to generate the quaternary structural proteome of E. coli MG1655 and demonstrate two use-cases involving large-scale mutant analysis and genome-scale modelling.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 97%
- Self-Supervised Deep-Learning Encodes High-Resolution Features of Protein Subcellular Localization 96%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
Similar papers in this journal
- Generalizable and scalable protein stability prediction with rewired protein generative models 96%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 96%
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 95%
Similar papers in this journal
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 96%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 95%
- AI-guided pipeline for protein-protein interaction drug discovery identifies a SARS-CoV-2 inhibitor 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.