Back

Identifying crosstalks among post-translational modifications in lung cancer proteomic data

Lai, S.; DAI, S.; Zhao, P.; Zhou, C.; Li, N.; Yu, W.

2024-08-08 bioinformatics
10.1101/2024.08.06.606765 bioRxiv
Show abstract

Post-translational modifications (PTMs) are pivotal in cellular regulations, and their crosstalk is related to various diseases such as cancer. Given the prevalence of PTM crosstalk within close amino acid ranges, identifying peptides with multiple PTMs is essential. However, this task is an NP-hard combinatorial problem with exponential complexity, posing significant challenges for existing analysis methods. Here, we introduce PIPI-C (PTM-Invariant Peptide Identification with a Combinatorial model), a novel search engine that addresses this challenge through a mixed-integer linear programming (MILP) model, thereby overcoming the limitations of existing approaches that struggle with high-order PTM combinations. Rigorous validation across diverse datasets confirms PIPI-Cs superior performance in detecting PTM crosstalks. When applied to over 72 million mass spectra of three human cancers--lung squamous cell carcinoma (LSCC), colorectal adenocarcinoma (COAD), and glioblastoma (GBM)--PIPI-C reveals significantly upregulated PTM crosstalks. In LSCC, 50% of 860 upregulated unique PTM site patterns (UPSPs) (when comparing cancer vs. normal samples) carried at least two PTMs, including literature-supported crosstalks such as di-methylation with trifluoroleucine substitution and amidation with proline-to-valine substitution. Similar findings in COAD and GBM highlight PIPI-Cs utility in uncovering cancer-relevant PTM crosstalk landscapes. Overall, PIPI-C provides a robust mathematical framework for decoding complex PTM patterns, advancing our understanding of PTM-driven cellular processes in diseases.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.