MoDPA: Inferring Modification-dependent Protein Associations from Uniformly Reprocessed Mass Spectrometry
Massignani, E.; Tichshenko, N.; Velghe, K.; Martens, L.
Show abstract
MotivationPost-translational modifications (PTMs) are key regulators of protein function and cellular processes; however, the overall principles of PTM co-regulation and crosstalk remain to be fully understood. A major challenge in large-scale PTM crosstalk studies is the scarcity of data, which hampers the ability to reproducibly detect proteome-wide associations between modification sites. ResultsWe present a new computational framework, named Modification-Dependent Protein Associations (MoDPA). To overcome the extreme sparsity and heterogeneity of PTM calls across experiments, MoDPA utilizes a variational autoencoder (VAE) to embed per-site detection profiles into a low-dimensional latent space that preserves covariation while denoising missing data. A PTM association network is constructed by correlating latent representations across experiments. Benchmarking against pulsed SILAC data shows that MoDPA can correctly capture the correlation between heavy- and light-labelled peptides. We apply MoDPA to a large sample of reprocessed public datasets and identify clusters of modified proteins involved in distinct biological pathways, suggesting that some modifications may preferentially regulate specific biological processes, but not others. In particular, we find clusters enriched in lysine acetylation, lysine methylation, and arginine deamidation (citrullination) sites, which contain proteins involved in protein synthesis, neuron development, and cellular senescence. Availability and ImplementationAll code is freely available on GitHub (https://github.com/CompOmics/MoDPAv1.0) and Zenodo (10.5281/zenodo.18310674). The resulting MoDPA-derived PTM network is available via the TabloidProteome website at https://iomics.ugent.be/tabloidproteome. Contactenrico.massignani@ugent.be, lennart.martens@ugent.be
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mokapot: Fast and flexible semi-supervised learning for peptide detection 97%
- Middle-down proteomics reveals dense sites of methylation and phosphorylation in arginine-rich RNA-binding proteins 96%
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 95%
Similar papers in this journal
- Integrated view and comparative analysis of baseline protein expression in mouse and rat tissues 94%
- Multienzyme deep learning models improve peptide de novo sequencing by mass spectrometry proteomics 94%
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 94%
Similar papers in this journal
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
- Proteoform identification using multiplexed top-down mass spectra 94%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 94%
Similar papers in this journal
- Transcriptome-informed reduction of protein databases: an analysis of how and when proteogenomics enhances eukaryotic proteomics 95%
- HUBMet: An integrative database and analytical platform for human blood metabolites and metabolite-protein associations 92%
- Protein length distribution is remarkably consistent across Life 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.