Back

Boosting metaproteomics identification rates and taxonomic specificity with MS2Rescore

Van Den Bossche, T.; Declercq, A.; Gabriels, R.; Holstein, T.; Mesuere, B.; Muth, T.; Verschaffelt, P.; Martens, L.

2025-02-22 bioinformatics
10.1101/2025.02.17.638783 bioRxiv
Show abstract

BackgroundMetaproteomics, the study of the collective proteome within microbial ecosystems, has gained increasing interest over the past decade. However, peptide identification rates in metaproteomics remain low compared to single-species proteomics. A key challenge is the identification sensitivity of current identification algorithms, which were primarily designed for single-species analyses. Addressing this, we evaluated the machine learning-driven MS{superscript 2}Rescore post-processing tool on multiple metaproteomics datasets from diverse microbial environments and benchmark studies. ResultsWe demonstrate that machine learning-driven rescoring outperforms traditional metaproteomics identification workflows. It significantly increases peptide identification rates compared to Sage, which itself already implements basic rescoring. Moreover, it enables lowering the false discovery rate (FDR) to 0.1% with minimal to no sensitivity loss, a substantial improvement over the 1% or 5% FDR thresholds commonly used in metaproteomics, in turn leading to greater confidence in downstream taxonomic annotation. ConclusionsOur findings show that MS{superscript 2}Rescore substantially improves peptide identification sensitivity as well as specificity in metaproteomics, and delivers improved confidence in taxonomic annotation. This advancement results in a more reliable downstream taxonomic analysis, reinforcing the potential of machine learning-based rescoring in metaproteomics research.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.