Tesorai Search: Large pretrained model boosts identifications in mass spectrometry proteomics without the need for Percolator.
Burq, M.; Stepec, D.; Restrepo, J.; Zbontar, J.; Urazbakhtin, S.; Crampton, B.; Tiwary, S.; Miao, M.; Cox, J.; Cimermancic, P.
Show abstract
The original mass spectrometry search engines used simple algorithms for peptide identification. Recent tools improved accuracy by adding several extra components such as fragment ion intensities or retention times prediction and training target-decoy classifiers on-the-fly, leading to sometimes inconsistent results. Our study explores the impact of replacing those extra components with a deep-learning pretrained model that directly learns the complex relationship between the full spectra and associated peptide sequence, without using decoys. This simplified workflow has fewer parameters to tweak, making it easier to use and perform robustly on data from instruments and use-cases never seen during training. Surprisingly, our approach consistently identifies more peptides than FragPipe, PEAKS, and Proteome Discoverer (12%, 9%, and 21% more, respectively, across a range of datasets). Tesorai Search is also fast - 250 immunopeptidomics searches in 45 minutes - and free for academics, available as a webserver at console.tesorai.com.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Proteoform identification using multiplexed top-down mass spectra 97%
- Performance Characteristics of Zeno Trap Scanning DIA for Sensitive and Quantitative Proteomics at High Throughput 95%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 95%
Similar papers in this journal
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 98%
- MSFragger-DDA+ Enhances Peptide Identification Sensitivity with Full Isolation Window Search 97%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 96%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 97%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 96%
- MS2AI: Automated repurposing of public peptide LC-MS data for machine learning applications 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.