De novo peptide databases enable protein-based stable isotope probing of microbial communities with up to species-level resolution
Klaes, S.; White, C.; Alvarez-Cohen, L.; Adrian, L.; Ding, C.
Show abstract
BackgroundProtein-based stable isotope probing (Protein-SIP) is a powerful approach that can directly link individual taxa to activity and substrate assimilation, elucidating metabolic pathways and trophic relationships within microbial communities. In Protein-SIP, peptides and corresponding taxa are identified by database matching, making database quality crucial for accurate analyses. For samples with unknown community composition, Protein-SIP typically employs either unrestricted reference databases or metagenome-derived databases. While (meta)genome-derived databases represent the gold standard, they may be incomplete and are typically resource-intensive to generate. In contrast, unrestricted reference databases can inflate the search space and require complex post-processing. ResultsHere, we explore the feasibility of using de novo peptide sequencing to construct peptide databases directly from mass spectrometry raw data. We then use the mass spectrometric data from labeled cultures to quantify isotope incorporation into specific peptides. We benchmark our approach against the canonical approach in which a sample-matching (meta)genome-derived protein sequence database is used on three different datasets: 1) a proteome analysis from a defined microbial community containing 13C-labeled E. coli cells, 2) time-course data of an anammox-dominated continuous reactor after feeding with 13C-labeled bicarbonate, and 3) a model of the human distal gut simulating a high-protein and high-fiber diet cultivated in either 2H2O or H218O. Our results show that de novo peptide databases are applicable to different isotopes, detecting similar amounts of labeled peptides compared to sample-matching (meta)genome-derived databases, and also identify labeled peptides missed by this canonical approach. Furthermore, we show that peptide-centric Protein-SIP allows up to species-specific resolution and enables the assessment of activity related to individual biological processes. Finally, we provide access to our modular Python pipeline to assist the construction of de novo peptide databases and subsequent peptide-centric Protein-SIP data analysis (https://git.ufz.de/meb/denovo-sip). ConclusionsDe novo peptide databases enable Protein-SIP of microbial communities without prior knowledge of the composition and can be used complementarily to (meta)genome-derived databases or as a standalone alternative in exploratory or resource-limited settings.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The E. coli PeptideAtlas Build: Characterizing the observed Escherichia coli pan-proteome and its post-translational modifications 96%
- Biological Function Assignment Across Taxonomic Levels in Mass-Spectrometry-Based Metaproteomics via a Modified Expectation Maximization Algorithm 96%
- The use of hybrid data-dependent and -independent acquisition spectral libraries empower dual-proteome profiling 96%
Similar papers in this journal
- Ultra-sensitive Protein-SIP to quantify activity and substrate uptake in microbiomes with stable isotopes 98%
- Critical Assessment of MetaProteome Investigation 2 (CAMPI-2): Multi-laboratory assessment of sample processing methods to stabilize fecal microbiome for functional analysis 97%
- MetaboDirect: An Analytical Pipeline for the processing of FTICR-MS-based Metabolomics Data 95%
Similar papers in this journal
- Comparative Performance of Scribe and Database Search Engines in Metaproteomic Profiling of a Ground-Truth Microbiome Dataset 96%
- A systematic evaluation of yeast sample preparation protocols for spectral identifications, proteome coverage and post-isolation modifications 94%
- ibaqpy: A scalable Python package for baseline quantification in proteomics leveraging SDRF metadata 94%
Similar papers in this journal
- Critical Assessment of Metaproteome Investigation (CAMPI): a Multi-Lab Comparison of Established Workflows 96%
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 95%
- Proteome allocation is linked to transcriptional regulation through a modularized transcriptome 94%
Similar papers in this journal
- Lost and found: re-searching and re-scoring proteomics data aids the discovery of bacterial proteins and improves proteome coverage 95%
- Absolute proteome quantification in the gas-fermenting acetogen Clostridium autoethanogenum 94%
- Mass spectrometry imaging of natural carbonyl products directly from agar-based microbial interactions using 4-APEBA derivatization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.