How to train a post-processor for tandem mass spectrometry proteomics database search while maintaining control of the false discovery rate
Freestone, J. A.; Käll, L.; Noble, W. S.; Keich, U.
Show abstract
Decoy-based methods are a popular choice for the statistical validation of peptide detections in tandem mass spectrometry proteomics data. Such methods can achieve a substantial boost in statistical power when coupled with post-processors such as Percolator that use auxiliary features to learn a better-discriminating scoring function. However, we recently showed that Percolator can struggle to control the false discovery rate (FDR) when reporting the list of discovered peptides. To address this problem, we introduce Percolator-RESET, which is an adaptation of our recently developed RESET meta-procedure to the peptide detection problem. Specifically, Percolator-RESET fuses Percolators iterative SVM training procedure with RESETs general framework of determining the list of reported discoveries in a target-decoy competition setup, where each putative discovery is augmented with a list of relevant features. Percolator-RESET operates in both a standard single-decoy mode and a two-decoy mode, the latter requiring the generation of two decoys per target. We demonstrate that Percolator-RESET controls the FDR in both modes, both theoretically and empirically, while typically reporting only a marginally smaller number of discoveries than Percolator in single-decoy mode. The two-decoy mode is marginally more powerful than both Percolator and the single-decoy mode and exhibits less variability than the latter.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Semi-supervised Bayesian integration of multiple spatial proteomics datasets 93%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 93%
- Designing diverse and high-performance proteins with a large language model in the loop 93%
Similar papers in this journal
- Assessment of false discovery rate control in tandem mass spectrometry analysis using entrapment 97%
- Unsupervised removal of systematic background noise from droplet-based single-cell experiments using CellBender 92%
- PTM-Mamba: A PTM-Aware Protein Language Model with Bidirectional Gated Mamba Blocks 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.