MetaNovo: a probabilistic approach to peptide and polymorphism discovery in complex mass spectrometry datasets
Potgieter, M. G.; Nel, A. J.; Tabb, D. L.; Fortuin, S.; Garnett, S.; Wendoh, J. M.; Blackburn, J.; Mulder, N.
Show abstract
BackgroundMicrobiome research is providing important new insights into the metabolic interactions of complex microbial ecosystems involved in fields as diverse as the pathogenesis of human diseases, agriculture and climate change. Poor correlations typically observed between RNA and protein expression datasets make it hard to accurately infer microbial protein synthesis from metagenomic data. Additionally, mass spectrometry-based metaproteomic analyses typically rely on focussed search libraries based on prior knowledge for protein identification that may not represent all the proteins present in a set of samples. Metagenomic 16S rRNA sequencing will only target the bacterial component, while whole genome sequencing is at best an indirect measure of expressed proteomes. We describe a novel approach, MetaNovo, that combines existing open-source software tools to perform scalable de novo sequence tag matching with a novel algorithm for probabilistic optimization of the entire UniProt knowledgebase to create tailored databases for target-decoy searches directly at the proteome level, enabling analyses without prior expectation of sample composition or metagenomic data generation, and compatible with standard downstream analysis pipelines. ResultsWe compared MetaNovo to published results from the MetaPro-IQ pipeline on 8 human mucosal-luminal interface samples, with comparable numbers of peptide and protein identifications, many shared peptide sequences and a similar bacterial taxonomic distribution compared to that found using a matched metagenome database - but simultaneously identified many more non-bacterial peptides than the previous approaches. MetaNovo was also benchmarked on samples of known microbial composition against matched metagenomic and whole genomic database workflows, yielding many more MS/MS identifications for the expected taxa, with improved taxonomic representation, while also highlighting previously described genome sequencing quality concerns for one of the organisms, and identifying a known sample contaminant without prior expectation. ConclusionsBy estimating taxonomic and peptide level information directly on microbiome samples from tandem mass spectrometry data, MetaNovo enables the simultaneous identification of peptides from all domains of life in metaproteome samples, bypassing the need for curated sequence search databases. We show that the MetaNovo approach to mass spectrometry metaproteomics is more accurate than current gold standard approaches of tailored or matched genomic database searches, can identify sample contaminants without prior expectation and yields insights into previously unidentified metaproteomic signals, building on the potential for complex mass spectrometry metaproteomic data to speak for itself. The pipeline source code is available on GitHub1 and documentation is provided to run the software as a singularity-compatible docker image available from the Docker Hub2.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A systematic evaluation of yeast sample preparation protocols for spectral identifications, proteome coverage and post-isolation modifications 95%
- Identification of plasma proteins associated with oesophageal cancer chemotherapeutic treatment outcomes using SWATH-MS 95%
- Proteomics study of colorectal cancer and adenomatous polyps identifies TFR1, SAHH, and HV307 as potential biomarkers for screening 95%
Similar papers in this journal
- Candida albicans: a comprehensive view of the proteome 96%
- PAMPA: a software for peptide markers and taxonomic identification for ZooMS samples in Archaeology and Paleontology 96%
- Massive proteogenomic reanalysis of publicly available proteomic datasets of human tissues in search for protein recoding via adenosine-to-inosine RNA editing 96%
Similar papers in this journal
- Leveraging immonium ions for identifying and targeting acyl-lysine modifications in proteomic datasets 95%
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 94%
Similar papers in this journal
- Proteomic and metabolomic signatures associated with the immune response in healthy individuals immunized with an inactivated SARS-CoV-2 vaccine 93%
- Impact of inflammatory preconditioning on murine microglial proteome response induced by focal ischemic brain injury 93%
- Accurate MHC Motif Deconvolution of Immunopeptidomics Data Reveals a Significant Contribution of DRB3, 4 and 5 to the Total DR Immunopeptidome 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.