Back

Journal of Proteomics

Elsevier BV

All preprints, ranked by how well they match Journal of Proteomics's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
MetaNovo: a probabilistic approach to peptide and polymorphism discovery in complex mass spectrometry datasets

Potgieter, M. G.; Nel, A. J.; Tabb, D. L.; Fortuin, S.; Garnett, S.; Wendoh, J. M.; Blackburn, J.; Mulder, N.

2019-07-11 bioinformatics 10.1101/605550 medRxiv
Top 0.1%
42.9%
Show abstract

BackgroundMicrobiome research is providing important new insights into the metabolic interactions of complex microbial ecosystems involved in fields as diverse as the pathogenesis of human diseases, agriculture and climate change. Poor correlations typically observed between RNA and protein expression datasets make it hard to accurately infer microbial protein synthesis from metagenomic data. Additionally, mass spectrometry-based metaproteomic analyses typically rely on focussed search libraries based on prior knowledge for protein identification that may not represent all the proteins present in a set of samples. Metagenomic 16S rRNA sequencing will only target the bacterial component, while whole genome sequencing is at best an indirect measure of expressed proteomes. We describe a novel approach, MetaNovo, that combines existing open-source software tools to perform scalable de novo sequence tag matching with a novel algorithm for probabilistic optimization of the entire UniProt knowledgebase to create tailored databases for target-decoy searches directly at the proteome level, enabling analyses without prior expectation of sample composition or metagenomic data generation, and compatible with standard downstream analysis pipelines. ResultsWe compared MetaNovo to published results from the MetaPro-IQ pipeline on 8 human mucosal-luminal interface samples, with comparable numbers of peptide and protein identifications, many shared peptide sequences and a similar bacterial taxonomic distribution compared to that found using a matched metagenome database - but simultaneously identified many more non-bacterial peptides than the previous approaches. MetaNovo was also benchmarked on samples of known microbial composition against matched metagenomic and whole genomic database workflows, yielding many more MS/MS identifications for the expected taxa, with improved taxonomic representation, while also highlighting previously described genome sequencing quality concerns for one of the organisms, and identifying a known sample contaminant without prior expectation. ConclusionsBy estimating taxonomic and peptide level information directly on microbiome samples from tandem mass spectrometry data, MetaNovo enables the simultaneous identification of peptides from all domains of life in metaproteome samples, bypassing the need for curated sequence search databases. We show that the MetaNovo approach to mass spectrometry metaproteomics is more accurate than current gold standard approaches of tailored or matched genomic database searches, can identify sample contaminants without prior expectation and yields insights into previously unidentified metaproteomic signals, building on the potential for complex mass spectrometry metaproteomic data to speak for itself. The pipeline source code is available on GitHub1 and documentation is provided to run the software as a singularity-compatible docker image available from the Docker Hub2.

2
Proteomic Dynamics and Heat Stress Response in Arabidopsis Seedlings

Fan, K.-T.; Xu, Y.

2024-04-30 biochemistry 10.1101/2024.04.30.591943 medRxiv
Top 0.1%
31.9%
Show abstract

Global warming poses a grave threat to plant survival, adversely affecting growth and agricultural productivity. To develop thermotolerant crops, a profound comprehension of plant responses to heat stress at the molecular level is imperative. Leveraging a novel fusion of 15N-stable isotope labeling and the ProteinTurnover algorithm, we meticulously investigated proteome dynamics in Arabidopsis thaliana seedlings subjected to moderate heat stress (30{degrees}C). This innovative approach facilitated a comprehensive analysis of proteomic changes across diverse cellular fractions. Our study unveiled significant turnover rate alterations in 571 proteins, with a median increase of 1.4-fold, indicative of accelerated protein dynamics under heat stress. Notably, root soluble proteins exhibited more subdued changes, suggesting tissue-specific adaptations. Moreover, we observed noteworthy turnover variations in proteins associated with redox signaling, stress response, and metabolism, underscoring the complexity of the response network. Conversely, proteins involved in carbohydrate metabolism and mitochondrial ATP synthesis displayed minimal turnover changes, signifying their stability. This exhaustive examination sheds light on the proteomic adjustments of Arabidopsis seedlings to moderate heat stress, elucidating the delicate balance between proteome stability and adaptability. These findings significantly augment our understanding of plant thermal resilience and offer crucial insights for the development of crops endowed with enhanced thermotolerance.

3
PTMOverlay: A Proteomic Tool to Visualize Post-Translational Modifications Across Evolution

Krieger, C.; Everton, Z.; You, Y.; Lewis, B.; Bank, T.; Burnet, M. C.; Williams, S.; Walukiewicz, H.; Rao, C.; Wolfe, A.; Payne, S. H.; Nakayasu, E. S.

2026-02-06 systems biology 10.64898/2026.02.03.703592 medRxiv
Top 0.1%
30.9%
Show abstract

Evolutionary conservation has been considered a hallmark of essential basic functions in cells. Therefore, the study of evolutionarily conserved post-translational modifications (PTMs) can provide insight into their role in protein function. In this context, mass spectrometry can identify and quantify thousands of PTM sites. However, a major bottleneck lies in analyzing the large amounts of data collected by the mass spectrometer. Here we address the need for a protein sequence alignment tool for multiple PTMs across several species. We developed a tool named PTMOverlay that takes peptide identification output files and overlays PTM sites onto multiple protein sequence alignments. Examining 31 bacteria isolates, we combined their protein sequences with select PTM types, including acetylation, phosphorylation, monomethylation, dimethylation, and trimethylation. The tool revealed a variety of conserved modification sites on the bacterial central carbon metabolism. Further structural analysis revealed possible interactions between methylated arginine and lysine residues with phosphothreonine/serine sites on the homodimer interface of enolase. Overall, this tool can parse large amounts of mass spectrometry data and allows for more informed and efficient selection of sites for future studies of protein function.

4
Lysis buffer selection guidance for mass spectrometry-based global proteomics including studies on the intersection of signal transduction and metabolism

Helm, B.; Hansen, P.; Lai, L.; Schwarzmuller, L.; Clas, S. M.; Richter, A.; Ruwolt, M.; Liu, F.; Frey, D.; DAlessandro, L. A.; Lehmann, W. D.; Schilling, M.; Helm, D.; Fiedler, D.; Klingmuller, U.

2024-02-21 systems biology 10.1101/2024.02.19.580971 medRxiv
Top 0.1%
29.7%
Show abstract

Prerequisite for a successful proteomics experiment is a high-quality lysis of the sample of interest, resulting in a large number of identified proteins as well as a high coverage of protein sequences. Therefore, the choice of suitable lysis conditions is crucial. Many buffers were previously employed in proteomics studies, yet a comprehensive comparison of lysate preparation conditions was so far missing. In this study, we compared the efficiency of four commonly used lysis buffers, containing the agents NP40, SDS, urea or GdnHCl, in four different types of biological samples (suspension and adherent cell lines, primary mouse cells and mouse liver tissue). After liquid chromatography-mass spectrometry (LC-MS) measurement and MaxQuant analysis, we compared chromatograms, intensities, number of identified proteins and the localization of the identified proteins. Overall, SDS emerged as the most reliable reagent, ensuring stable performance and reproducibility across diverse samples. Furthermore, our data advocated for a dual-sample lysis approach, including that the resulting pellet is lysed again after the initial lysis with a urea lysis buffer and subsequently both lysates are combined for a single LC-MS run to maximize the proteome coverage. However, none of the investigated lysis buffers proved to be superior in every category, indicating that the lysis buffer of choice depends on the proteins of interest and on the biological question. Further, we demonstrated with our systematic studies the establishment of conditions that allows to perform global proteomics and affinity purification-based interactome characterization from the same lysate. In sum our results provide guidance for the best-suited lysis buffer for mass spectrometry-based proteomics depending on the question of interest.

5
Peptidomics Mapping of Proteolysis Highlights Triple Activation of Sprouted Seeds by Germination, Homogenisation and Species Mixture.

Bera, I.; Fernandez-Diaz, R.; O Sullivan, M.; Jacquir, J.-C.; Scaife, C.; Litovskich, G.; Wynne, K.; Shields, D.

2026-01-09 bioinformatics 10.64898/2026.01.08.698447 medRxiv
Top 0.1%
27.7%
Show abstract

We investigated how seed proteolysis was enhanced by germination, by subsequent homogenisation (disrupting sprout compartments), and by co-incubation of homogenates from different species. Mass spectrometry of released peptides tracked proteolytic signatures from chickpea, lentil, mung and broccoli proteins, in soaked seeds, in sprouted seeds, and after sprout homogenisation followed by incubation alone or in mixture with other sprouts. The proteolytic signatures differed markedly among the four species, and in the different treatment conditions. After homogenisation, legumain-like cleavage (after asparagine) increased in lentil, and proline-rich peptides increased in broccoli. For co-incubated homogenised sprouts, each species homogenate significantly contributed 6 to 57% of proteolytic patterns in peptides of other species, with chickpea and broccoli homogenates notably releasing metabolic protein peptides from mung and lentil. Thus, germination, homogenisation and homogenate species mixtures can each contribute to proteolysis of seed peptides, potentially increasing digestibility and reducing allergenicity. HIGHLIGHTSO_LIProteolysis motifs in soaked seeds and sprouts very diverse among species C_LIO_LISeed germination proteolysis altered by homogenisation C_LIO_LISeed germination proteolysis altered by co-incubation of different species C_LIO_LIFoods based on homogenised sprout mixtures may release more digestible peptides C_LI

6
From the lung to the muscle: Systemic insights from an integrative MultiOmics analysis of harbour porpoises in poor respiratory health

Dönmez, E. M.; Siebels, B.; Drotleff, B.; Nissen, P.; Derous, D.; Fabrizius, A.; Siebert, U.

2026-03-31 systems biology 10.64898/2026.03.28.714973 medRxiv
Top 0.1%
27.5%
Show abstract

Harbour porpoises (Phocoena phocoena) in the North and Baltic Seas are increasingly impacted by anthropogenic pressures, including underwater noise, fisheries and pollution. These pressures correlate with declining population health, particularly affecting the respiratory system. Growing pathological lesions, partly resulting from high prevalence of parasitic infestations and subsequent diseases, can impair tissue function and oxygen supply to distant end-organs. In this study, we applied an integrative MultiOmics approach (proteomics, metabolomics, lipidomics) to analyse the lungs and muscles of 12 wild harbour porpoises with compromised respiratory health. Our aim was to identify dysregulated biological pathways across omics layers to advance insights into adaptive physiological responses and to define disease-associated molecular signatures that could assist health assessments. Our analysis revealed pronounced immune system and antioxidative responses in the lungs and muscles, indicated by enhanced immunoglobulins, plasmalogens and glutathione-related proteins. In the lungs, high cardiolipin levels and reduced collagen suggest impaired tissue structure and function, while tissue maintenance processes were elevated in the muscle. Both tissues exhibited metabolic alterations suggestive of energetic imbalance, including increased purine metabolism in the lung and decreased lipid metabolism in the muscle. Several dysregulated molecules were shared across tissues, pointing to pathophysiological effects. The proposed disease-associated molecular signatures included the protein SLC25A4, the metabolite O-phosphoethanolamine and the lipid TG O-16:0_16:0_20:4 for the lung, and the protein SPEG, the metabolite pipecolic acid, and the lipid BMP 18:1_22:6 in the muscle. Our findings elucidate the complexity of molecular mechanisms linking anthropogenic and environmental stressors with vulnerability and resilience in a marine sentinel species. Furthermore, this study highlights the potential of integrative omics to define disease-related marker panels, thereby supporting ongoing and future health monitoring and conservation efforts.

7
Post-translational modification fidelity of recombinant human lactopontin expressed in Kluyveromyces lactis

Excell, J.; Giardina, A.; Sakamoto-Rablah, E.; Royle, K.; Nunn, D.

2026-05-12 synthetic biology 10.64898/2026.05.12.724256 medRxiv
Top 0.1%
22.9%
Show abstract

Recombinant human lactopontin (rhLPN), an equivalent of human milk lactopontin, is of increasing interest for human nutrition applications due to its roles in mineral binding, gastrointestinal function and immune modulation. These properties depend strongly on post-translational modifications, particularly phosphorylation and glycosylation. Here, we report the production of rhLPN in Kluyveromyces lactis at laboratory and pilot scale and present a comprehensive molecular comparison with native human lactopontin (nhLPN) isolated from human milk. Mass spectrometry-based peptide mapping confirmed the primary structure and identified extensive phosphorylation, consistent with the native protein. Middle-up analyses demonstrated closely matched phosphoform distributions between rhLPN and nhLPN, while glycosylation profiling revealed a defined population of low-complexity O-glycoforms localized to the N-terminus. Functional assessment demonstrated substantially greater iron binding by phosphorylated rhLPN compared with dephosphorylated and non-phosphorylated forms. Similar phosphorylation-dependent behaviour was observed for bovine lactopontin, supporting a conserved role for phosphorylation in mineral interaction. Across five 750 L pilot scale batches, both phosphorylation and glycoform distributions were highly consistent, indicating robust process reproducibility. Together, these results demonstrate that rhLPN produced in K. lactis recapitulates key structural and functional attributes of nhLPN, supporting its suitability as a scalable ingredient for nutrition applications.

8
Spatial differentiation of proteome in cervical cancer tissues using Imaging Mass Spectrometry

Baliyan, A.; Mohapatra, I.; Paul, S.; Safiriyu, A. A.; Mondal, S. K.; Mishra, D.; Mandal, A. K.

2025-04-18 cancer biology 10.1101/2025.04.12.648495 medRxiv
Top 0.1%
22.8%
Show abstract

Cervical cancer, which is the fourth most common gynaecological cancer across the globe, has a poorly understood molecular pathogenesis and etiology. Current methods of diagnosis are based on cytology, histology and presence of Human Papillomavirus (HPV). Shortcomings of these methods lie with poor quality of smears which is very common in Papanicolaou (pap) smears, limited sensitivity in terms of early detection, dependence on presence of HPV, etc. Due to its high sensitivity and non-targeted approach, Matrix Assisted Laser Desorption Ionization based Imaging Mass Spectrometry (MALDI-IMS) might be advantageous to understand the pathogenesis, molecular mechanism, precise identification of surgical margins and identification of novel biomarkers. Since it can also identify proteins in the extracellular matrix, it is especially beneficial for the tissue types with sparse cells and excessive extracellular matrix. Although tissue proteome profiling for cervical cancer were reported, the heterogeneous distribution of proteins across cervical cancer tissues havent been explored. In this study, we employed a non-targeted MALDI-IMS based approach to profile the spatial distribution of proteins within cervical cancer tissues. We observed overexpression of Keratin 5 and Prelamin A/C in the region of cervical cancer tissues which were categorically labelled with cancerous morphology using histopathological examination. Both these proteins have been earlier associated with progression and aggressiveness of other cancers like breast and prostate cancers. However, no such reports are available for cervical cancer. Further studies are required on a large dataset to validate and quantitate these proteins as biomarkers for early diagnosis and prognosis of cervical cancer.

9
Detecting predicted cancer-testis antigens in proteomics datasets of healthy and tumoral samples

Machado, K. C. T.; Fiuza, T. D. S.; De Souza, S. J.; De Souza, G. A.

2024-06-09 bioinformatics 10.1101/2024.06.08.597624 medRxiv
Top 0.1%
22.6%
Show abstract

Biomarkers are molecular markers found in clinical samples which may aid disease diagnosis or prognosis. High-throughput techniques allow prospecting for such signature molecules by comparing gene expression between normal and sick cells. Cancer-testis antigens (CTAs) are promising candidates for cancer biomarkers due to their limited expression to the testis in normal conditions versus their aberrant expression in various tumors. CTAs are routinely identified by transcriptomics, but a comprehensive characterization of their protein levels in different tissues is still necessary. Mass spectrometry-based proteomics allows the characterization of many cellular types and the production of large amounts of data while computational tools allow the comparison of multiple datasets, and together those may corroborate insights obtained at the transcriptomic level. Here a computational meta-analysis explores the CTAs protein abundance in the proteomic layer of healthy and tumor tissues. The combined datasets present the expression patterns of 17,200 unique proteins, including 241 known CTAs previously described at the transcriptomic level. Those were further ranked as significantly enriched in tumor tissues (22 proteins), exclusive to tumor tissues (42 proteins) or abundant in healthy tissues (32 proteins). This analysis illustrates the possibilities for tumor proteome characterization and the consequent identification of biomarker candidates and/or therapeutic targets.

10
DIA-PASEF Proteomic Profiling Reveals MpkA-Dependent Iron Stress Responses and Siderophore Biosynthesis in Aspergillus nidulans

Lee, J.; West, O. K.; Huso, W.; Doan, A. G.; Gray, K. J.; Edwards, H.; Tran, J. T.; Carman, D. R.; Betenbaugh, M.; Srivastava, R.; Harris, S.; Marten, M. R.

2025-07-06 systems biology 10.1101/2025.07.03.662993 medRxiv
Top 0.1%
22.5%
Show abstract

Filamentous faungi play essential roles in biotechnology as producers of valuable bioproducts and, conversely, as opportunistic pathogens. Aspergillus nidulans is a widely used model organism for fungal genetics and cell biology; however, comprehensive proteomic references for this species remain limited. In this study, we applied a data-independent acquisition-parallel accumulation serial fragmentation (DIA-PASEF) approach to enable efficient and in-depth proteome profiling of A. nidulans. Leveraging ion mobility-based ion cloud information from DDA-PASEF experiments, we developed DIA-PASEF methods that identified 3,904 proteins across biological triplicates grown in rich medium. Compared to prior studies, this represents an increase in protein identifications of more than 140% and was achieved with more than five-fold reduction in analysis time. We employed this newly developed DIA-PASEF methodology to conduct both proteomic and phosphoproteomic analyses under iron-depleted conditions in an MpkA protein-kinase deficient mutant ({Delta}mpkA). The {Delta}mpkA strain exhibited expression of approximately 500 additional proteins and occupancy of over 1,800 additional phosphosites relative to a control. Differentially expressed and phosphorylated proteins increased by more than an order of magnitude in the {Delta}mpkA mutant across both iron-replete and iron-deplete conditions. Gene Ontology (GO) enrichment analysis revealed broader and distinct biological processes under iron-depleted conditions, highlighting adaptive responses specific to iron limitation and MAPK pathway disruption. This work establishes a high-coverage proteomic resource for A. nidulans and provides novel insights into fungal stress responses and signaling network perturbation. Importantly, high-throughput proteomic profiling reveals that limited iron availability and MAPK pathway disruption increases siderophore biosynthesis.

11
A comparison of four proteomics software for hair proteome analyses

Mukonyora, M.

2026-04-20 bioinformatics 10.64898/2026.04.17.719199 medRxiv
Top 0.1%
20.2%
Show abstract

1.1Hair has applications in biomarker discovery and forensics, yet the influence of proteomics software tools on hair proteome characterisation remains underexplored. This study compares four bottom-up proteomics workflows (MaxQuant, FragPipe, MetaMorpheus, and SearchGUI/PeptideShaker). Publicly available hair proteomes were analysed following extraction with 1-dodecyl-3-methylimidazolium chloride (DMC), sodium dodecanoate (SDD), sodium dodecyl sulfate (SDS), and urea. Data were acquired on Orbitrap-based DDA platforms. Peptide identification, protein inference, functional annotation, physicochemical properties, and label-free quantification (LFQ) were evaluated. Peptide-level performance differed across tools. MS-GF+ and FragPipe identified the most unique peptides, while X!Tandem reported the fewest. Protein inference showed a dissociation from peptide-level results. MetaMorpheus reported the highest number of protein groups despite only the third highest peptide counts. FragPipe and MaxQuant followed, while PeptideShaker consistently inferred the fewest proteins. Protein-level concordance was low, with only 30.3% overlap across tools and extraction methods. These differences extended to downstream analyses. Functional enrichment showed moderate concordance (38.25% overlap). Physicochemical profiles varied, with MetaMorpheus identifying more hydrophobic proteomes and PeptideShaker more hydrophilic profiles. At the quantitative level, reproducibility depended on extraction buffer. SDS and urea showed lower variability (CV =< 0.025), while DMC and SDD showed higher variability (up to 0.10). Absolute LFQ intensities and differential expression outputs varied across tools despite moderate to strong correlation (r = 0.77 to 0.93). Overall, software choice influences proteome coverage, physicochemical profiles, and quantitative outcomes. Relative trends were partially conserved, but magnitude and significance varied. These findings support careful method selection and multi-tool validation in hair proteomics

12
SPROUTS_DB: an implemented database of contaminants for extracellular vesicle proteomics studies

Pittala, M. G. G.; Leggio, L.; Paterno, G.; Giusto, E.; Civiero, L.; Cunsolo, V.; Vivarelli, S.; Di Francesco, A.; Alpi, E.; Saletti, R.; Iraci, N.

2025-05-21 cell biology 10.1101/2025.05.20.655024 medRxiv
Top 0.1%
19.7%
Show abstract

BackgroundCurrent proteomics techniques allow rapid identification and quantification of proteins within any given biological source. In particular, nanoUHPLC/High-Resolution nanoESI-MS/MS enables the characterization of proteins in complex biological samples due to its high sensitivity, accuracy, and scalability. However, LC-MS/MS proteomics might still be susceptible to laboratory and sample-associated contaminants, which can significantly compromise the quality and reliability of data. Therefore, an accurate identification and annotation of such contaminants is crucial for the development of robust proteomics databases and spectral-libraries related search engines. This approach is of special interest in the field of secretome and extracellular vesicles (EVs), membrane-enclosed nanostructures that contain a variety of proteins crucial for cell-to-cell communication and translational applications. ResultsWhen working in ex vivo/in vitro settings, proteins from fetal bovine serum (FBS), commonly employed in standard cell culture media, may interfere with the proteome analysis. To address this issue, we conceived and designed SPROUTS_DB, Serum Protein Repository Of Unwanted Target(ed) Sequences DataBase, a dedicated resource to catalog serum-derived contaminants. Starting from media supplemented with EV-depleted FBS, we simulated cell growth conditions - in the absence of cells - followed by ultracentrifugation. LC-MS/MS analysis of these samples resulted in the identification of a novel set of 1,288 contaminant proteins, which has been deposited in the ProteomeXchange repository (identifier PXD044137). SPROUTS_DB contains primarily soluble proteins, mainly related to the Gene Ontology categories Extracellular Region and Extracellular Space, in line with the nature of the starting sample. In contrast, only a small fraction of the contaminants is classified as membrane-associated proteins, supporting the limited vesicle contamination in the complete medium, due to the use of EV-depleted FBS. Of note, we demonstrated that SPROUTS_DB outperforms existing contaminants databases, ensuring that only peptide spectra relevant to the examined sample are retained and identified as true positive data. ConclusionsConsidering that even proteins from phylogenetically distant organisms share extensive stretches of sequences, SPROUTS_DB is designed to discern contaminants from real sample proteins of interest, minimizing false positive identifications. To the best of our knowledge, SPROUTS_DB is the most updated database of contaminants useful for proteomics investigations of cellular secretomes and EV-containing samples.

13
Comparative Evaluation of DDA and DIA Based Proteomic Workflows in Beryllium Related Lung Disease

Weise, D. O.; Gupta, K.; Griffin, T. J.; Jagtap, P. D.; Mroz, M. M.; Wagner, R.; Macaluso, J. D.; Mehta, S.; Maier, L. A.; Li, L.; Vestal, B. E.; Bhargava, M.

2026-06-09 systems biology 10.64898/2026.06.04.730108 medRxiv
Top 0.1%
19.6%
Show abstract

We compared traditional data-dependent acquisition mass spectrometry (DDA-MS) with the increasingly adopted data-independent acquisition (DIA-MS) to evaluate their relative utility for large-scale quantitative biofluid proteomics of lung compartments, specifically paired bronchoalveolar lavage (BAL) cells and bronchoalveolar lavage fluid (BALF). Using beryllium-related granulomatous lung disease as a focused model, we analyzed BALF and BAL cells from beryllium-sensitized (BeS) individuals using both acquisition strategies to assess proteome depth, quantitative completeness, and analytical robustness. In BAL cells, 5,640 proteins were identified by DDA-MS and 5,227 by DIA-MS; however, DIA-MS yielded markedly improved quantitative completeness, with 5,178 proteins ([~]99%) quantified across all samples compared with 3,539 ([~]63%) quantified by DDA-MS. While 3,397 proteins were quantified by both methods, DIA-MS uniquely quantified 1,781 lower-abundance proteins. Proteins identified by both DIA and DDA-MS approaches revealed pathways associated with granulomatous inflammation, including Toll-like receptor, clathrin-mediated endocytosis, sirtuin, and C-type lectin receptor signaling, whereas DIA-MS resolved additional pathways, such as the complement cascade, coagulation system, and JAK/IL-6-type cytokine signaling. In BALF, although more proteins were identified by DDA-MS than by DIA-MS (2,069 vs 1,742), DIA-MS achieved greater quantitative completeness, with 1,695 proteins quantified across all samples compared with 1,050 using DDA-MS, underscoring its suitability for biomarker-oriented analyses in lung fluid compartments. Together, these results support DIA-MS as a robust and sensitive platform for quantitative lung proteomics and discovery of disease-relevant protein signatures.

14
A meta-analysis of rice phosphoproteomics data to understand variation in cell signalling across the rice pan-genome

Ramsbottom, K. A.; Prakash, A. A.; Perez-Riverol, Y.; Camacho, O. M.; Sun, Z.; Kundu, D.; Bowler-Barnett, E.; Martin, M.; Fan, J.; Chebotarov, D.; McNally, K.; Deutsch, E. W.; Vizcaino, J. A.; Jones, A. R.

2023-11-17 bioinformatics 10.1101/2023.11.17.567512 medRxiv
Top 0.1%
19.2%
Show abstract

Phosphorylation is the most studied post-translational modification, and has multiple biological functions. In this study, we have re-analysed publicly available mass spectrometry proteomics datasets enriched for phosphopeptides from Asian rice (Oryza sativa). In total we identified 15,522 phosphosites on serine, threonine and tyrosine residues on rice proteins. We identified sequence motifs for phosphosites, and link motifs to enrichment of different biological processes, indicating different downstream regulation likely caused by different kinase groups. We cross-referenced phosphosites against the rice 3,000 genomes, to identify single amino acid variations (SAAVs) within or proximal to phosphosites that could cause loss of a site in a given rice variety. The data was clustered to identify groups of sites with similar patterns across rice family groups, for example those highly conserved in Japonica, but mostly absent in Aus type rice varieties - known to have different responses to drought. These resources can assist rice researchers to discover alleles with significantly different functional effects across rice varieties. The data has been loaded into UniProt Knowledge-Base - enabling researchers to visualise sites alongside other data on rice proteins e.g. structural models from AlphaFold2, PeptideAtlas and the PRIDE database - enabling visualisation of source evidence, including scores and supporting mass spectra.

15
Identification of astrocytomas through serum protein fingerprint using MALDI-TOF MS and machine learning

Lazari, L. C.; Silva, J. M.; Donado, P. R. S.; Shinjo, S. M. O.; Fernandes, L. R.; Ieva, A. D.; Palmisano, G.; Marie, S. K. N.

2025-02-27 systems biology 10.1101/2025.02.21.639567 medRxiv
Top 0.1%
19.2%
Show abstract

Gliomas account for most brain malignancies, with astrocytomas being the most common subtype. Among these, glioblastoma (GBM) stands out as the most aggressive form, exhibiting a median survival time of just 15 months despite intensive therapy. Current diagnostic practices rely on magnetic resonance imaging (MRI) and histopathological analysis, which often necessitate invasive surgical sampling. This underscores the need for minimally invasive diagnostic tools capable of characterizing glioma progression and guiding treatment strategies. Advances in glioma classification have integrated histological and molecular markers, notably IDH1 mutations, which are prognostically significant, particularly in low-grade gliomas and in the previously defined "secondary GBM" (IDH-mutant astrocytoma grade 4). This study aimed to explore the potential of serum proteomics as a non-invasive diagnostic tool using MALDI-TOF mass spectrometry (MS) combined with machine learning techniques. We analyzed serum samples from 269 patients, employing machine learning models to differentiate between healthy individuals and astrocytoma patients. The MALDI-TOF MS approach achieved a balanced accuracy of 94.5% in distinguishing GBM patients from healthy controls. However, it showed limited efficacy in classifying tumor grades or determining IDH1 mutational status. Further investigation using bottom-up proteomics by GeLC-MS/MS identified potential biomarkers, such as transthyretin, previously associated with high-grade gliomas. These findings highlight the promise of MALDI-TOF MS in identifying serum-based biomarkers for astrocytoma diagnosis. While the results are promising, further validation in independent cohorts is essential to assess the clinical utility of these biomarkers for non-invasive glioma diagnostics and patient monitoring.

16
Proteomics for cultivated meat: the importance of Analytical Standardization

Palma, J.; Leblanc, C. C.; Kusters, R.; Kamgang Nzekoue, A. F.

2026-03-25 systems biology 10.64898/2026.03.23.713501 medRxiv
Top 0.1%
19.2%
Show abstract

Cultivated meat production requires robust and validated analytical methods for comprehensive characterization. While transcriptomics-based approaches establish the foundational profile of molecular analysis, proteomics provides additional resolution that further enhances scientific certainty in both product development and safety characterization. However, the industry adoption of proteomics is currently hindered by technical complexity and a critical lack of analytical standardization, which leads to significant workflow-dependent variations in proteome coverage. To address this gap, we investigated the influence of key workflow steps (digestion, cleanup, LC-MS conditions) on the proteome profile of cultivated duck biomass. We compared five bottom-up sample preparation protocols - two traditional in-solution options (urea and SDC-based protocols), two device-based approaches (PreOmics iST and EasyPep kits), and an innovative protocol (SPEED), and demonstrated that device-based protocols offered the highest peptide yield and proteome coverage. However, optimization allowed cost-effective in-solution methods to achieve comparable performance. Specifically, an optimal digestion time of 3 hours at 37{degrees}C and the use of polymer-based desalting columns significantly enhanced protein identification ([~]4500 - 5000 IDs). Moreover, data independent acquisition (DIA) provided deeper proteome coverage than data dependent acquisition (DDA) with higher precision ([~]6500 vs 5000 IDs). The validated Standard Operating Procedures presented here establish a standardized framework for bulk bottom-up proteomics in cultivated meat, facilitating the generation of reliable and comparable data required for robust multi-omics characterization. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/713501v1_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@5b61b8org.highwire.dtl.DTLVardef@16c7e65org.highwire.dtl.DTLVardef@1de21d2org.highwire.dtl.DTLVardef@7e984a_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIComplexity and non-standardization limit MS-proteomics use in cultivated meat (CM). C_LIO_LICM protein profile varies with sample prep, LC-MS, and data processing pipeline. C_LIO_LIDevice-based and optimized cost-effective protocols offer a high proteome coverage. C_LIO_LIProteomics can complement transcriptomics for a comprehensive CM characterization. C_LIO_LIProposed standardized methods ensure reliable data for future regulatory submissions. C_LI

17
Temporal phosphoproteomics reveals rapid restoration of kinase signaling by Glycyrrhiza glabra in a rotenone-induced Parkinson disease model

Narayana, V. K.; Karthikkeyan, G.; Najar, M. A.; Pervaje, R.; T S, K. P.; Modi, P. K.

2026-06-18 neuroscience 10.64898/2026.06.15.732239 medRxiv
Top 0.1%
19.2%
Show abstract

Parkinsons disease is a progressive neurodegenerative disorder associated with mitochondrial dysfunction, oxidative stress, impaired autophagy, and dysregulated cellular signaling pathways. Although Glycyrrhiza glabra has been reported to exhibit neuroprotective properties, the early phosphorylation-mediated signaling mechanisms underlying its protective effects remain poorly understood. In this study, we employed a Tandem Mass Tag (TMT)-based temporal quantitative phosphoproteomic approach to investigate early signaling events associated with Glycyrrhiza glabra-mediated neuroprotection in a rotenone-induced in vitro PD model. Differentiated IMR-32 neuronal cells were treated with rotenone alone or in combination with Glycyrrhiza glabra extract, and phosphoproteomic alterations were analyzed at 2, 5, 15, and 30 minutes using liquid chromatography coupled with tandem mass spectrometer. Temporal phosphoproteomic analysis identified 6,424 phosphopeptides corresponding to 2,368 phosphoproteins and 5,468 phosphorylation sites. Comparative analysis revealed extensive phosphorylation rewiring induced by rotenone and restoration of several dysregulated phosphorylation events following Glycyrrhiza glabra co-treatment. More than 130 phosphoproteins and multiple kinase-associated signaling pathways were dynamically regulated across the temporal conditions. Kinase enrichment analysis identified restoration of several critical kinases, including AKT1, MTOR, MAPK1/3, PRKACA, PRKCD, and GSK3A/B, which are associated with neuronal survival, stress adaptation, and autophagy. Integrated pathway and kinase-substrate interaction analyses further revealed enrichment of AMPK signaling, FOXO signaling, receptor tyrosine kinase signaling, RNA processing, and cell-cycle regulatory pathways. Notably, several spliceosome-associated phosphoproteins demonstrated dynamic phosphorylation changes during the early neuroprotective response. Collectively, this study provides a detailed temporal phosphoproteomic landscape of early signaling events associated with Glycyrrhiza glabra-mediated neuroprotection and highlights kinase-driven signaling pathways that may represent potential therapeutic targets in Parkinsons disease.

18
Borrelia PeptideAtlas: A proteome resource of common Borrelia burgdorferi isolates for Lyme research

Reddy, P. J.; Sun, Z.; Wippel, H. H.; Baxter, D. H.; Swearingen, K. E.; Shteynberg, D. D.; Midha, M. K.; Caimano, M. J.; Strle, K.; Choi, Y.; Chan, A. P.; Schork, N. J.; Moritz, R. L.

2023-06-16 biochemistry 10.1101/2023.06.16.545244 medRxiv
Top 0.1%
18.7%
Show abstract

Lyme disease, caused by an infection with the spirochete Borrelia burgdorferi, is the most common vector-borne disease in North America. B. burgdorferi strains harbor extensive genomic and proteomic variability and further comparison is key to understanding the spirochetes infectivity and biological impacts of identified sequence variants. To achieve this goal, both transcript and mass spectrometry (MS)-based proteomics was applied to assemble peptide datasets of laboratory strains B31, MM1, B31-ML23, infective isolates B31-5A4, B31-A3, and 297, and other public datasets, to provide a publicly available Borrelia PeptideAtlas (http://www.peptideatlas.org/builds/borrelia/). Included is information on total proteome, secretome, and membrane proteome of these B. burgdorferi strains. Proteomic data collected from 35 different experiment datasets, with a total of 855 mass spectrometry runs, identified 76,936 distinct peptides at a 0.1% peptide false-discovery-rate, which map to 1,221 canonical proteins (924 core canonical and 297 noncore canonical) and covers 86% of the total base B31 proteome. The diverse proteomic information from multiple isolates with credible data presented by the Borrelia PeptideAtlas can be useful to pinpoint potential protein targets which are common to infective isolates and may be key in the infection process.

19
Proteomics analysis in urinary bladder cancer patients identifies urinary SOD2 as a predictive marker of recurrence

Kumari, N.; Biswal, S. C.; Chaudhary, S.; Malalkar, D.; Dubey, U. S.; Vasudevae, P.; Kumar, A.; Saxena, S.; NANDA, R.; Agrawal, U.

2021-12-14 oncology 10.1101/2021.12.13.21267125 medRxiv
Top 0.1%
18.5%
Show abstract

Early non-invasive detection of tumor is an urgent clinical need for managing urothelial bladder cancer. Cystoscopy and cytology are the current standards for diagnosis of recurrence, but are limited by low sensitivity. Quantitative proteomics tool was employed to identify important deregulated molecules in bladder cancer tissues and validated using Western blot and immunohistochemistry analysis. A set of 1137 proteins were identified in four paired bladder cancer patients. Among these, 64 proteins were deregulated in all cases among which 9 were commonly up-regulated. The Ingenuity Pathway Analysis (IPA) generated top 11 Networks in which three commonly upregulated (SERPING1, SOD2 and HSPB6) proteins were involved and selected for further validation. Tissue expression of SOD2, SERPING1 and HSPB6 monitored in an independent sample set (n=18) by immuno-histochemical analysis showed similar profile. Western blot analysis of these proteins in urine of bladder cancer (n=26) and healthy subjects (n=10) showed a specificity and sensitivity of >80% for SOD2 and so was selected for further validation in a separate set (n=150) by ELISA. Significant elevation in urinary SOD2 level was found in urothelial bladder cancer patients compared to healthy controls and in recurrent cases compared to primary (p-value<0.001). Kaplan Meier survival analysis showed urinary SOD2 concentration >2,100 pg/ml was significantly associated with poorer survival.Cumulative survival of patient with low SOD2 concentration was 34.4% compared to 18.9% in patient with high SOD2 at 24 months (p=0.025). The study identifies SOD2 as a non-invasive biomarker which may help to extend the period between cystoscopies during follow-up. SignificanceCystoscopy is an invasive and painful method commonly used for diagnosis of urothelial bladder cancer. Non-invasive methods having high specificity and sensitivity to monitor the patients for recurrence are unavailable. Our study reveals significantly higher SOD2 level in drug naive and reoccurring bladder cancer tissues, and similar profile was observed in the parallel urine samples. Hence, SOD2 seems to be a useful biomarker of recurrent urothelial bladder cancer and predict the survival of patients.

20
Glycosylated proteins identified for the first time in the alga Euglena gracilis

O'Neill, E. C.

2021-10-28 biochemistry 10.1101/2021.10.28.466288 medRxiv
Top 0.1%
18.5%
Show abstract

Protein glycosylation, and in particular N-linked glycans, is a hall mark of Eukaryotic cells and has been well studied in mammalian cells and parasites. However, little research has been conducted to investigate the conservation and variation of protein glycosylation pathways in other eukaryotic organisms. Euglena gracilis is an industrially important microalga, used in the production of biofuels and nutritional supplements. It is evolutionarily highly divergent from green algae and more related to Kinetoplastid pathogens. It was recently shown that E. gracilis possesses the machinery for producing a range of protein glycosylations and make simple N-glycans, but the modified proteins were not identified. This study identifies the glycosylated proteins, including transporters, extra cellular proteases and those involved in cell surface signalling. Notably, many of the most highly expressed and glycosylated proteins are not related to any known sequences and are therefore likely to be involved in important novel functions in Euglena.