Back

Journal of Proteome Research

American Chemical Society (ACS)

All preprints, ranked by how well they match Journal of Proteome Research's content profile, based on 234 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
One-Pot Quantitation of Mycobacterial N-terminal Protein Acetylation Peptidoforms and Proteome

Hu, D. D.; Weaver, S. D.; Jones, B. S.; Champion, P. A.; Champion, M. M.

2025-03-31 biochemistry 10.1101/2025.03.30.646241 medRxiv
Top 0.1%
84.6%
Show abstract

Isotopic labeling of proteins for quantitative proteomics is a popular technique to increase sample throughput and provide improved accuracy and precision for relative or absolute quantitation between samples. Derivatives of this technique are used to label protein and peptide N-termini for selective enrichment and analysis. We previously reported on a method to enrich and quantify protein N-terminal acetylation in model-system and pathogenic mycobacteria. Significant recent advancements in silica filter-based protein digestion have improved identification of proteins in bottom-up proteomics. However, these are not yet compatible with existing methods which detect and quantify protein termini and N-terminal modifications. Here, we present a one-pot method (OnePotNTA) that incorporates silica filter digestion with protein N-terminal labeling and subsequent quantitation. This technique achieves high-density coverage of the N-terminome and obviates the need to enrich N-terminal peptides prior to analysis. OnePotN TA identified 54% of the canonical proteome from whole-cell lysates of Mycobacterium marinum and eliminated biases in peptides identified by forgoing enrichment. This significantly reduced sample preparation time by [≥]5-fold and preserves protein-level abundance measurements (LFQ) from the same injections and analysis. Analysis of a mutant strain of M. marinum lacking Emp1 ({Delta}emp1)-an N-terminal Acetyltransferase required for efficient pathogenesis-identified 37 putative substrates of this enzyme. Additionally, analysis of the remaining peptides identified at least 34 proteins with alternate true N-termini, distinct from the canonical genome.

2
MARLOWE: Taxonomic Characterization of Unknown Samples for Forensics Using De Novo Peptide Identification

Jenson, S. C.; Chu, F.; Alves, G.; Ogurtsov, A. Y.; Barente, A. S.; Crockett, D. L.; Lamar, N. C.; Merkley, E. D.; Yu, Y.-K.; Jarman, K. H.

2025-06-02 bioinformatics 10.1101/2024.09.30.615220 medRxiv
Top 0.1%
81.3%
Show abstract

We present a computational tool, MARLOWE, for source organism characterization of unknown, forensic biological samples. The intent of MARLOWE is to address a gap in applying proteomics data analysis to forensic applications. MARLOWE produces a list of potential source organisms given confident peptide tags derived from de novo peptide sequencing and a statistical approach to assign peptides to organisms in a probabilistic manner, based on a broad sequence database. In this way, the algorithm assumes no a priori knowledge of potential sources, and the probabilistic way peptides are taxonomically assigned and then scored enables results to be unbiased (within the constraints of the sequence database). In a proof-of-concept study, we examined MARLOWEs performance on two datasets, the Biodiversity dataset and the Bacillus cereus superspecies dataset. Not only did MARLOWE demonstrate successful characterization to true contributors in single source and binary mixtures in the Biodiversity dataset, but also provided sufficient specificity to distinguish species within a bacterial superspecies group. We also compared MARLOWEs results to those of MiCId, a leading microbial identification/characterization tool based on proteomics database search. Comparison of the two tools using 225 mass spectrometry data files yielded comparable performance, with slightly higher accuracy and specificity for MiCId. At the species level, MARLOWE achieved a specificity of 91.4% at 5% FDR. These results suggest that MARLOWE is suitable for candidate- or lead-generation identification of single-organism and binary samples that can generate forensic leads and aid in selecting appropriate follow-on analyses in a forensic context.

3
Joint Protein Inference Analysis with PyProteinInference Elucidates Biological Understanding of Tandem Mass Spectrometry Data

Hinkle, T. B.; Bakalarski, C. E.

2024-10-20 bioinformatics 10.1101/2024.10.17.618892 medRxiv
Top 0.1%
77.3%
Show abstract

Selection and application of protein inference algorithms can have a significant impact on the data output from tandem mass spectrometry (MS/MS) experiments, yet its use is often an afterthought in proteomics research due to the inability to apply different inference algorithms in existing analysis systems today. PyProteinInference provides a comprehensive suite of tools to guide researchers through the application of multiple inference algorithms and computation of protein-level, set-based false discovery rates (FDR) from tandem mass spectrometry (MS/MS) data using a unified interface. Here, we describe the software and its application to a K562 whole-cell lysate as well as in a CRAF affinity-purification mass spectrometry experiment to demonstrate its utility in facilitating conclusions about underlying biological mechanisms in proteomic data.

4
Mining Mass Spectra for Peptide Facts

Lemieux, S.; Zumer, J.

2023-11-01 bioinformatics 10.1101/2023.10.27.564468 medRxiv
Top 0.1%
76.2%
Show abstract

The current mainstream software for peptide-centric tandem mass spectrometry data analysis can be categorized as either database-driven, which rely on a library of mass spectra to identify the peptide associated with novel query spectra, or de novo sequencing-based, which aim to find the entire peptide sequence by relying only on the query mass spectrum. While the first paradigm currently produces state-of-the-art results in peptide identification tasks, it does not inherently make use of information present in the query mass spectrum itself to refine identifications. Meanwhile, de novo approaches attempt to solve a complex problem in one go, without any search space constraints in the general case, leading to comparatively poor results. In this paper, we decompose the de novo problem into putatively easier subproblems, and we show that peptide identification rates of database-driven methods may be improved in terms of peptide identification rate by solving one such subsproblem without requiring a solution for the complete de novo task. We demonstrate this using a de novo peptide length prediction task as the chosen subproblem. As a first prototype, we show that a deep learning-based length prediction model increases peptide identification rates in the ProteomeTools dataset as part of an Pepid-based identification pipeline. Using the predicted information to better rank the candidates, we show that combining ideas from the two paradigms produces clear benefits in this setting. We propose that the next generation of peptide-centric tandem mass spectrometry identification methods should combine elements of these paradigms by mining facts "de novo; about the peptide represented in a spectrum, while simultaneously limiting the search space with a peptide candidates database.

5
Universal toolset for mass spectrometric analysis of intracellular peptidome and small protein fraction

Kote, S.; Pirog, A.; Faktor, J.; Dziadosz, A.; Marek-Trzonkowska, N.

2024-12-12 cancer biology 10.1101/2024.12.06.627199 medRxiv
Top 0.1%
75.8%
Show abstract

The analysis of native intracellular peptidome has gained significant attention in recent years. However, there is still a need for more knowledge regarding various sample preparation methods that facilitate efficient and reproducible recovery of peptides, which can then be analyzed using quantitative liquid chromatography-mass spectrometry. A similar situation exists in small proteome research, typically defined as polypeptides with masses of less than 100 amino acids, often too long for easy identification without enzymatic digestion. In this context, we describe a set of methods that involve simple denaturation and solid-phase extraction of polypeptides, applicable for isolating short intracellular polypeptides within the desired length range. Our work demonstrates the efficiency and reproducibility of these methods for quantitative analysis of the peptidome in mammalian cells. Additionally, we investigated the flexibility of adjusting the mass range through ultrafiltration. We have shown that these methods can be adapted for highly efficient enrichment and fractionation of small proteins, resulting in polypeptide isolates suitable for tryptic digestion and intact protein analysis. Moreover, we describe the use of freely available computational tools that can effectively manage the analysis of the resulting data. The research presented here will benefit the global scientific community in both fundamental (protein turnover, proteolytic processing, non-canonical open reading frames, etc.) and applied sciences (bioactive/neuro peptide discovery, precision medicine, vaccines, etc.), and other areas that could benefit from selective analysis of short native polypeptides.

6
ProteoDUDes: Taxonomic profiling for metaproteomics with false positive reduction

Schiebenhoefer, H.; Muth, T.; Fuchs, S.; Renard, B. Y.

2026-07-02 bioinformatics 10.64898/2026.06.29.734936 medRxiv
Top 0.1%
75.2%
Show abstract

Metaproteomics is the investigation of the protein composition of multi-organism samples. While metagenomics answers the question which organisms are present in a sample, metaproteomics additionally answers the question which organisms are active. State-of-the-art tools for annotating proteomic data with taxonomic information (e.g. Unipept, DIAMOND) do not control the false taxonomic identification rate, which can lead to incorrect results and thus incorrect interpretations, as we demonstrate with examples. ProteoDUDes processes the results from popular sequence annotation tools so that the proportion of true identifications in the result is at as high as or higher than in the the compared tools. We evaluate ProteoDUDes on simulated data and experimental mock community data. Our results indicate that ProteoDUDes has the same error rate as other tools on simulated data and half the error rate on the experimental mock community data. This allows more accurate statements to be made about which organisms are functionally active in a complex sample. ProteoDUDes is open-source and available at https://github.com/pirovc/dudes.

7
Proteogenomics analysis of human tissues using pangenomes

Wang, D.; Bouwmeester, R.; Zheng, P.; Dai, C.; Puente, A. S.; Shu, K.; Bai, M.; Umer, H. M.; Perez-Riverol, Y.

2024-05-28 bioinformatics 10.1101/2024.05.24.595489 medRxiv
Top 0.1%
74.2%
Show abstract

The genomics landscape is evolving with the emergence of pangenomes, challenging the conventional single-reference genome model. The new human pangenome reference provides an extra dimension by incorporating variations observed in different human populations. However, the increasing use of pangenomes in human reference databases poses challenges for proteomics, which currently relies on UniProt canonical/isoform-based reference proteomics. Including more variant information in human proteomes, such as small and long open reading frames and pseudogenes, prompts the development of complex proteogenomics pipelines for analysis and validation. This study explores the advantages of pangenomes, particularly the human reference pangenome, on proteomics, and large-scale proteogenomics studies. We reanalyze two large human tissue datasets using the quantms workflow to identify novel peptides and variant proteins from the pangenome samples. Using three search engines SAGE, COMET, and MSGF+ followed by Percolator we analyzed 91,833,481 MS/MS spectra from more than 30 normal human tissues. We developed a robust deep-learning framework to validate the novel peptides based on DeepLC, MS2PIP and pyspectrumAI. The results yielded 170142 novel peptide spectrum matches, 4991 novel peptide sequences, and 3921 single amino acid variants, corresponding to 2367 genes across five population groups, demonstrating the effectiveness of our proteogenomics approach using the recent pangenome references.

8
Unipept in 2024: Expanding metaproteomics analysis with support for missed cleavages, semi-tryptic and non-tryptic peptides

Vande Moortele, T.; Devlaminck, B.; Van de Vyver, S.; Van Den Bossche, T.; Martens, L.; Dawyndt, P.; Mesuere, B.; Verschaffelt, P.

2024-11-27 bioinformatics 10.1101/2024.09.26.615136 medRxiv
Top 0.1%
73.8%
Show abstract

Unipept, a pioneering software tool in metaproteomics, has significantly advanced the analysis of complex ecosystems by facilitating both taxonomic and functional insights from environmental samples. From the onset, Unipepts capabilities focused on tryptic peptides, utilizing the predictability and consistency of trypsin digestion to efficiently construct a protein reference database. However, the evolving landscape of proteomics and emerging fields like immunopeptidomics necessitate a more versatile approach that extends beyond the analysis of tryptic peptides. In this article, we present a significant update to the underlying index structure of Unipept, which is now powered by a Sparse Suffix Array index. This advancement enables the analysis of semi-tryptic peptides, peptides with missed cleavages, and non-tryptic peptides such as those encountered in other research fields such as immunopeptidomics (e.g. MHC- and HLA-peptides). This new index benefits all tools in the Unipept ecosystem such as the web application, desktop tool, API and command line interface. A benchmark study highlights significantly improved performance in handling missed cleavages, preserving the same level of accuracy. For TOC Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/615136v2_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@5b8fe2org.highwire.dtl.DTLVardef@1435321org.highwire.dtl.DTLVardef@106a568org.highwire.dtl.DTLVardef@15563e2_HPS_FORMAT_FIGEXP M_FIG C_FIG

9
DIGEST: An online tool for designing of multiple reaction monitoring assays

gautam, p.; Singh, P.; Mrinal, ; Bhaskar, A.; Sacher, S.; Dagar, Y.; Basak, T.; Sengupta, S.; Ray, A.

2023-11-28 bioinformatics 10.1101/2023.11.27.568790 medRxiv
Top 0.1%
73.7%
Show abstract

Targeted proteomics using multiple reaction monitoring (MRM) assays enables fast and sensitive detection of a preselected set of target peptides. This technique utilizes the specificity of precursors to product transitions for quantitative analysis of multiple proteins in a single sample. The success of an MRM experiment depends on the selection of transitions however, given the existing resources, accurately predicting signal intensity of peptides and their fragmentation patterns ab initio is challenging task. We present an alternative for rapid design of MRM transitions for proteomics research: DIGEST. Our method predicts the b and y ions with +1 and +2 charge produced in a collision cell of a mass spectrometer from peptides of multiple proteotypically digested proteins. Additionally, by using the existing knowledge of the fundamental rules for designing transitions, the tool provides optimal MRM transitions, negating the need to undertake prior "discovery" MS studies. We demonstrate that our algorithm is directed toward the selection of MRM precursor and product-ions pairs, and can avoid the pitfalls of interference due to cross-contamination of samples by selecting ion combinations that uniquely map to target peptides. Comparison with SRMAtlas showed that DIGEST successfully predicted the peptide and production pairs in the majority of cases. We believe that DIGEST will facilitate rapid design of MRM assays with increased specificity, reducing the overall time required to design an MRM assay for routine mass-spectrometry. DIGEST is available as a web-based tool at https://digest.raylab.iiitd.edu.in/

10
Proteogenomic analysis of the differential stability of cardiac protein isoforms

Juber, M.; Binti, S.; Pandi, B.; Lau, E.; Lam, M. P. Y.

2026-07-24 biochemistry 10.64898/2026.07.23.740363 medRxiv
Top 0.1%
72.9%
Show abstract

Alternative splicing is an important regulatory layer in gene expression, but knowledge on the isoform protein molecules continue to lag their canonical counterparts. An open question is whether alternative protein isoforms feature different half-life than the canonical counterpart, which could indicate differential usage and functional diversification. Here we combined a proteogenomics approach with heavy water-based protein turnover analysis to survey 24 pairs of canonical-alternative protein isoforms in the mouse heart. The results provide a reference on their numerical half-life and also reveal widespread differences in isoform stability.

11
In silico approach to accelerate the development of mass spectrometry-based proteomics methods for detection of viral proteins: Application to COVID-19

Jenkins, C.; Orsburn, B.

2020-03-10 biochemistry 10.1101/2020.03.08.980383 medRxiv
Top 0.1%
72.2%
Show abstract

We describe a method for rapid in silico selection of diagnostic peptides from newly described viral pathogens and applied this approach to SARS-CoV-2/COVID-19. This approach is multi-tiered, beginning with compiling the theoretical protein sequences from genomic derived data. In the case of SARS-CoV-2 we begin with 496 peptides that would be produced by proteolytic digestion of the viral proteins. To eliminate peptides that would cause cross-reactivity and false positives we remove peptides from consideration that have sequence homology or similar chemical characteristics using a progressively larger database of background peptides. Using this pipeline, we can remove 47 peptides from consideration as diagnostic due to the presence of peptides derived from the human proteome. To address the complexity of the human microbiome, we describe a method to create a database of all proteins of relevant abundance in the saliva microbiome. By utilizing a protein-based approach to the microbiome we can more accurately identify peptides that will be problematic in COVID-19 studies which removes 12 peptides from consideration. To identify diagnostic peptides, another 7 peptides are flagged for removal following comparison to the proteome backgrounds of viral and bacterial pathogens of similar clinical presentation. By aligning the protein sequences of SARS-CoV-2 field isolates deposited to date we can identify peptides for removal due to their presence in highly variable regions that may lead to false negatives as the pathogen evolves. We provide maps of these regions and highlight 3 peptides that should be avoided as potential diagnostic or vaccine targets. Finally, we leverage publicly deposited proteomics data from human cells infected with SARS-CoV-2, as well as a second study with the closely related MERS-CoV to identify the two proteins of highest abundance in human infections. The resulting final list contains the 24 peptides most unique and diagnostic of SARS-CoV-2 infections. These peptides represent the best targets for the development of antibodies are clinical diagnostics. To demonstrate one application of this we model peptide fragmentation using a deep learning tool to rapidly generate targeted LCMS assays and data processing method for detecting CoVID-19 infected patient samples. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=156 HEIGHT=200 SRC="FIGDIR/small/980383v2_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@1d7fd4borg.highwire.dtl.DTLVardef@136563borg.highwire.dtl.DTLVardef@57641dorg.highwire.dtl.DTLVardef@16de9a4_HPS_FORMAT_FIGEXP M_FIG C_FIG

12
Peptide-to-protein data aggregation using Fisher's method improves target identification in chemical proteomics

Lyu, H.; Gharibi, H.; Meng, Z.; Sokolova, B.; Zhang, X.; Zubarev, R.

2026-02-04 bioinformatics 10.64898/2026.02.02.702201 medRxiv
Top 0.1%
71.0%
Show abstract

Protein-level statistical tests in proteomics aimed at obtaining p-value are conventionally made on protein abundances aggregated from peptide data. This integral approach overlooks peptide-level heterogeneity and ignores important information coded in individual peptide data, while protein p-value can also be obtained by Fishers method of combining peptide p-values using chi-square statistics. Here we test this latter approach across diverse chemical proteomics datasets based on assessments of protein expression, solubility and protease accessibility. Using the top four peptides ranked by their p-values consistently outperformed protein-level analysis and avoided biases introduced by inclusion of deviant peptides or imputation of missing peptide values. Fishers method provides a simple and robust strategy, improving identification of regulated/shifted proteins in diverse proteomics assays.

13
Improving power while controlling the false discovery rate when only a subset of peptides are relevant

Lin, A.; Plubell, D. L.; Keich, U.; Noble, W. S.

2020-10-21 bioinformatics 10.1101/2020.10.20.347278 medRxiv
Top 0.1%
70.1%
Show abstract

The standard proteomics database search strategy involves searching spectra against a peptide database and estimating the false discovery rate (FDR) of the resulting set of peptide-spectrum matches. One assumption of this protocol is that all the peptides in the database are relevant to the hypothesis being investigated. However, in settings where researchers are interested in a subset of peptides, alternative search and FDR control strategies are needed. Recently, two methods were proposed to address this problem: subset-search and all-sub. We show that both methods fail to control the FDR. For subset-search, this failure is due to the presence of "neighbor" peptides, which are defined as irrelevant peptides with a similar precursor mass and fragmentation spectrum as a relevant peptide. Not considering neighbors compromises the FDR estimate because a spectrum generated by an irrelevant peptide can incorrectly match well to a relevant peptide. Therefore, we have developed a new method, "filter then subsetneighbor search" (FSNS), that accounts for neighbor peptides. We show evidence that FSNS properly controls the FDR when neighbors are present and that FSNS outperforms group-FDR, the only other method able to control the FDR relative to a subset of relevant peptides.

14
Cloud-based DIA data analysis module for signal refinement improves accuracy and throughput of large datasets.

Christianson, K. E.; Jaffe, J. D.; Carr, S. A.; Vaca Jacome, A. S.

2021-07-14 bioinformatics 10.1101/2021.07.14.452243 medRxiv
Top 0.1%
69.5%
Show abstract

Data-independent acquisition (DIA) is a powerful mass spectrometry method that promises higher coverage, reproducibility, and throughput than traditional quantitative proteomics approaches. However, the complexity of DIA data caused by fragmentation of co-isolating peptides presents significant challenges for confident assignment of identity and quantity, information that is essential for deriving meaningful biological insight from the data. To overcome this problem, we previously developed Avant-garde, a tool for automated signal refinement of DIA and other targeted mass spectrometry data. AvG is designed to work alongside existing tools for peptide detection to address the reliability and quantitative suitability of signals extracted for the identified peptides. While its use is straightforward and offers efficient refinement for small datasets, the execution of AvG for large DIA datasets is time-consuming, especially if run with limited computational resources. To overcome these limitations, we present here an improved, cloud-based implementation of the AvG algorithm deployed on Terra, a user-friendly cloud-based platform for large-scale data analysis and sharing, as an accessible and standardized resource to the wider community.

15
Integrated View of Baseline Protein Expression in Human Tissues using public Data Independent Acquisition datasets

Prakash, A.; Collins, A.; Vilmovsky, L.; Fexova, S.; Vizcaino, J. A.; Jones, A. R.

2024-09-19 bioinformatics 10.1101/2024.09.16.613191 medRxiv
Top 0.1%
68.9%
Show abstract

The PRIDE database is the largest public data repository of mass spectrometry-based proteomics data and currently stores more than 40,000 datasets covering a wide range of organisms, experimental techniques and biological conditions. During the past few years, PRIDE has seen a significant increase in the amount of submitted Data-Independent Acquisition (DIA) proteomics datasets. This provides an excellent opportunity for large scale data reanalysis and reuse. We have reanalysed 15 public label-free DIA datasets across various healthy human tissues, to provide a state-of-the-art view of the human proteome in baseline conditions (without any perturbations). We computed baseline protein abundances and compared them across various tissues, samples and datasets. Our second aim was to compare protein abundances obtained here from the results of previous analyses using human baseline Data-Dependent Acquisition (DDA) datasets. We observed a good correlation across some tissues, especially in liver and colon but weak correlations were found in others, such as lung and pancreas. The reanalysed results including protein abundance values and curated metadata are made available to view and download from the resource Expression Atlas. For TOC Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=126 SRC="FIGDIR/small/613191v2_ufig1.gif" ALT="Figure 1"> View larger version (26K): org.highwire.dtl.DTLVardef@1414dforg.highwire.dtl.DTLVardef@66606aorg.highwire.dtl.DTLVardef@143f862org.highwire.dtl.DTLVardef@1682406_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
From variability to consensus: rescoring harmonizes peptide identification across diverse search engines and datasets

Winkelhardt, D.; Berres, S.; Uszkoreit, J.

2026-03-06 bioinformatics 10.64898/2026.03.04.709532 medRxiv
Top 0.1%
68.2%
Show abstract

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, datasets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four datasets acquired on different mass spectrometry platforms and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human datasets, it significantly affected identification rates on a metaproteomic dataset. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a slightly higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

17
Sensitive and specific spectral library searching with COSS and Percolator

Shiferaw, G. A.; Gabriels, R.; Vandermarliere, E.; Martens, L.; Volders, P.-J.

2021-04-11 bioinformatics 10.1101/2021.04.09.438700 medRxiv
Top 0.1%
66.7%
Show abstract

Maintaining high sensitivity while limiting false positives is a key challenge in peptide identification from mass spectrometry data. Here, we therefore investigate the effects of integrating the machine learning-based post-processor Percolator into our spectral library searching tool COSS. To evaluate the effects of this post-processing, we have used forty data sets from two different projects and have searched these against the NIST and MassIVE spectral libraries. The searching is carried out using two spectral library search tools, COSS and MSPepSearch with and without Percolator post-processing, and using sequence database search engine MS-GF+ as a baseline comparator. The addition of the Percolator rescoring step to COSS is effective and results in a substantial improvement in sensitivity and specificity of the identifications. COSS is freely available as open source under the permissive Apache2 license, and binaries and source code are found at https://github.com/compomics/COSS

18
Digging deeper into the immunopeptidome with TripleToolWF

Mayer, R. L.; Mechtler, K.

2026-06-14 biochemistry 10.64898/2026.06.11.731513 medRxiv
Top 0.1%
66.4%
Show abstract

While the field of immunopeptidomics has matured substantially over the last years, high input amounts of cellular or tissue material are still required to obtain a somewhat complete profile of the immunopeptidome. Here we present a simple platform termed TripleToolWF (derived from Triple Tool workflow) to increase the number of identified and quantified immunopeptides combining the outputs of three search engines such as PEAKS Online 12, Sequest HT with INFERYS rescoring and MSFragger. For assessing the false discovery rate (FDR) an entrapment approach is used. The platform improved peptide identifications by 6-14% and peptide quantitations by 11-25% compared to the best individual search engine for two independent, previously published, bacterial infection datasets. Peptides were mostly 9-12mers as expected and >90% of the obtained 9mers were predicted binders by the stringent majority voting approach of Immunolyser 2.0 which indicates high confidence of the identified immunopeptides. The FDR was monitored using dedicated entrapment searches against shuffled databases. The resulting entrapment FDR was assessed before and after result pooling and showed only a minor increase upon pooling compared to the worst individual search engine. It remained even below the target of 1% peptide FDR in 40% of the experiments. Compared to the original publications, the number of high confidence bacterial immunopeptides was drastically elevated by 53% and 2800% for the Listeria monocytogenes and Mycobacterium bovis BCG projects, respectively, when applying strict filters. Of these additional bacterial sequences, all 9mer sequences were predicted as binders by at least one of the prediction algorithms of Immunolyser 2.0 illustrating their actual HLA binding nature. TripleToolWF hence provides a simple tool to further increase the number of obtained sequences from MS-based immunopeptidomics experiments to facilitate a deeper view of the immunopeptidome for refined vaccine candidate prioritization.

19
An LLM-driven pipeline for proteomics-based detection and structural modeling of post-translational modifications

George, A.; Mejia-Rodriguez, D.; Li, X.; Rigor, P.; Cheung, M. S.; Bilbao, A.

2026-05-06 bioinformatics 10.64898/2026.05.01.722279 medRxiv
Top 0.1%
66.2%
Show abstract

Post-translational modifications (PTMs) on proteins dynamically regulate their functions and subsequently cellular physiology. Significant advances have been made in their detection and modeling: mass spectrometry-based proteomics has become the cornerstone for PTM detection in complex samples, while emerging structure-prediction frameworks enable modeling of PTM-dependent conformational changes. However, the biological significance of many PTMs remains largely unexplored, in part because integrated pipelines that bridge PTM detection with structural modeling remain limited. We present a generative AI-driven pipeline that integrates PTM detection with structural modeling of their effects on protein dynamics and interactions. The pipeline comprises two complementary tools: PTMdiscoverer and PTM-Psi. First, PTMdiscoverer leverages large language models to identify, annotate, and interpret candidate PTMs from open-search proteomics results, addressing limitations of conventional proteomics tools. Next, PTM-Psi models the structural, functional, and dynamic consequences of these spatially aware modifications on protein dynamics. These two components bridge PTM discovery with mechanistic interpretation at the structural level. We demonstrate our pipeline by using cyanobacterial proteomics data to study potential molecular mechanisms of redox-regulated "dark complex" formation in carbon metabolism, advancing our ability to interpret PTM-mediated regulation in microbial systems.

20
Evaluation of Parallel Accumulation-Serial Fragmentation methods for metaproteomics using a model microbiome

Shrestha, R.; Rajczewski, A. T.; Do, K.; Willetts, M.; Kleiner, M.; Griffin, T.; Jagtap, P. D.

2025-08-15 biochemistry 10.1101/2025.08.13.670166 medRxiv
Top 0.1%
66.2%
Show abstract

Mass spectrometry-based metaproteomics allows for the identification and quantification of thousands of proteins from clinical and environmental samples and is rapidly gaining importance in microbiome sciences. Metaproteomics researchers can measure taxonomic and functional abundances of microbiomes, shedding light on mechanistic details of microbiome interactions with their environment. However, metaproteomic analysis suffers from limited depth of coverage due to the presence of millions of peptides at lower abundance levels. Recent advances in data-independent acquisition mass spectrometry coupled with Parallel Accumulation-Serial Fragmentation (PASEF) technology offer improved depth of coverage. PASEF technology enables simultaneous accumulation of ions from multiple co-eluting peptides by combining ion mobility separation with dynamic quadrupole isolation, allowing efficient and selective fragmentation in a single scan. This boosts ion sampling efficiency and resolves overlapping signals with high sensitivity. In this study, we assessed proteome coverage, quantitative precision, and accuracy of Data-dependent acquisition (DDA) and Data-independent acquisition (DIA) methods coupled with the PASEF method. For this, we used a ground-truth mock community containing 28 species (30 strains) from all three domains of life and bacteriophages with a 400-fold dynamic range of organism abundance. Our results showed that diaPASEF demonstrated superior performance, identifying 168% more peptide precursors, 155% more peptides, and 66% more protein groups compared to ddaPASEF. Quantitative measurements showed improved precision with diaPASEF, with 26 out of 28 organisms exhibiting coefficient of variation values below 20%, compared to 24 organisms with ddaPASEF. Both ddaPASEF and diaPASEF methods accurately quantified the 22 most abundant organisms, while measurements of low-abundance bacteriophages showed significant deviation from expected values. Our findings demonstrate that diaPASEF provides enhanced depth of coverage and quantitative reliability for metaproteomics analysis, particularly beneficial for clinical and environmental microbiome studies where deeper functional characterization is essential. This study provides valuable benchmark data to facilitate the development of advanced bioinformatic methods for quantitative metaproteomics.