A proteogenomics workflow to uncover the world of small proteins in Staphylococcus aureus
Fuchs, S.; Kucklick, M.; Lehmann, E.; Beckmann, A.; Wilkens, M.; Kolte, B.; Mustafayeva, A.; Ludwig, T.; Diwo, M.; Wissing, J.; Jaensch, L.; Ahrens, C. H.; Ignatova, Z.; Engelmann, S.
Show abstract
Small proteins play diverse and essential roles in bacterial physiology and virulence. Despite their importance, automated genome annotation algorithms still cannot accurately annotate all respective small open reading frames (sORFs), as they usually provide insufficient sequence information for domain and homology searches, tend to be species specific and only a few experimentally validated examples are covered in standard proteomics studies. The accuracy and reliability of genome annotations, particularly for sORFs, can be significantly improved by integrating protein evidence from experimental approaches that enrich for small proteins. Here we present a highly optimized and flexible workflow for bacterial proteogenomics, which covers all steps from (i) creation of protein databases, (ii) database searches, (iii) peptide-to-genome mapping to (iv) result interpretation and whose automated execution is supported by two open source tools (SALT & Pepper). We used the workflow to identify high quality peptide spectrum matches (PSMs) for both annotated and unannotated small proteins ([≤] 100 aa; SP100) in Staphylococcus aureus Newman. Proteins isolated from cells at the exponential and stationary growth phase were digested with different endopeptidases (trypsin, Lys-C, AspN), the resulting peptides fractionated by gel-based and gel-free methods and measured with highly sensitive mass spectrometers. PSMs or sORF predictions from sORFfinder were stringently filtered allowing us to detect 185 soluble SP100, 69 of which were missing in the used genome annotation. Most interestingly, almost half of the identified SP100 were basic, suggesting a role in binding to more acidic molecules such as nucleic acids or phospholipids. In addition, phage-related functions were proposed for 30 SP100, based on the localization of their coding sequences in the genome.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extending the proteomic characterization of Candida albicans exposed to stress and apoptotic inducers through data-independent acquisition mass spectrometry 94%
- Proteomic and transcriptomic analysis of Microviridae {varphi}X174 infection reveals broad up-regulation of host membrane damage and heat shock responses 94%
- Lost and found: re-searching and re-scoring proteomics data aids the discovery of bacterial proteins and improves proteome coverage 94%
Similar papers in this journal
- A proteogenomic resource enabling integrated analysis of Listeria genotype-proteotype-phenotype relationships 95%
- Integrative metabolomics and proteomics allow the global intracellular characterization of Bacillus subtilis cells and spores. 95%
- Comparative Analysis of Lysine-Specific Peptidases for Optimizing Proteomics Workflows 95%
Similar papers in this journal
Similar papers in this journal
- Assigning a role for chemosensory signal transduction in Campylobacter jejuni biofilms using a combined omics approach 94%
- The post-translational modification landscape of commercial beers 94%
- Cryptic, solo acylhomoserine lactone synthase from predatory myxobacterium suggests beneficial contribution to prey quorum signaling 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.