Improved peptide search for identification of SUMO and sequence-based modifications, in MaxSBM
Lennartsson, C.; Kyriakidou, P.; Nielsen, M. L.; Olsen, J. V.; Cox, J.; Hendriks, I. A.
Show abstract
Post-translational modifications (PTMs), such as SUMOylation and ubiquitination, regulate key cellular processes by covalently attaching to lysine residues. While mass spectrometry allows site-specific identification of PTMs, most existing search engines are optimized for small, non-fragmenting modifications and struggle to detect large, fragmenting protein-based modifiers. We refer to these as Sequence-Based Modifiers (SBMs). To overcome this limitation, we developed an SBM-specific search strategy within MaxQuant that accounts for the fragmentation behavior of SBMs during peptide identification. Using publicly available datasets, we validated our approach for SUMO2/3. Our analysis identified distinct diagnostic features and characteristic mass shifts associated with SBM fragmentation, referred to in this study as d-ions (diagnostic ions) and p-ions. By leveraging these features, our method improved the identification of SUMOylated peptides from human cell lines by 13%, SUMOylation sites in mouse embryonic cells by 18%, and in mouse adipocytes by 25%. Our search method improved spectral annotation of SBMs by up to 9% increase in the median Andromeda score. Taken together, we highlight the potential of our SBM search to enhance the discovery of protein-based modifications. HighlightsO_LIDevelopment of a MaxQuant module tailored for identifying Sequence-Based Modifiers (SBMs), including SUMO2/3 C_LIO_LIIncorporation of SBM-specific fragmentation patterns into search algorithms C_LIO_LIEnhanced biological discovery through improved PTM identification from mass spectrometry datasets C_LI In BriefHere, we introduce MaxSBM, an optimized framework for interpreting complex sequence-based modifiers (SBMs), particularly SUMO, within MaxQuant. Our approach incorporates SBM-specific d- and p-ion series into peptide scoring and annotation. By extending the theoretical spectral space to include fragments bearing partial SUMO (or other SBM) peptide remnants, MaxSBM provides a more comprehensive spectral annotation which enhances peptide scoring, resulting in increased identification rates at a higher confidence. Beyond methodological refinement, we validated MaxSBM via reanalysis of several physiological SUMO datasets, ultimately unlocking new insights via mapping of previously obscured modification sites.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mokapot: Fast and flexible semi-supervised learning for peptide detection 97%
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 96%
- Middle-down proteomics reveals dense sites of methylation and phosphorylation in arginine-rich RNA-binding proteins 96%
Similar papers in this journal
- Proteoform identification using multiplexed top-down mass spectra 96%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 95%
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
Similar papers in this journal
- ibaqpy: A scalable Python package for baseline quantification in proteomics leveraging SDRF metadata 97%
- AA_stat: intelligent profiling of in vivo and in vitro modifications from open search results 96%
- In-gel protein digestion using acidic methanol produces a highly selective methylation of glutamic 1 acid residues. 96%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 95%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 95%
- AlphaMap: An open-source Python package for the visual annotation of proteomics data with sequence specific knowledge 94%