ProteoForge: An Imputation-Aware Framework for Differential Proteoform Discovery in Bottom-Up Proteomics
Ergin, E. K.; Conrrero, A.; Ferguson, K. M.; Lange, P. F.
Show abstract
The human genome contains approximately 20,000 protein-coding genes. However, millions of diverse protein variants, called proteoforms, exist. Despite originating from the same gene, proteoforms often have distinct biological roles. In bottom-up proteomics, the aggregation of peptide measurements into protein-level quantities often obscures this information. Existing methods for proteoform deconvolution are limited by their handling of missing data, which can introduce significant bias. To address this we developed ProteoForge, which builds on an imputation-aware statistical model to identify and group co-varying peptides into quantitatively differential proteoforms (dPFs). Benchmarking against existing deconvolution methods demonstrated that ProteoForge provides high accuracy and stability in datasets with high rates of missing values, complex experimental designs, or varying signal strengths. Application of ProteoForge to proteomics data from lung cancer cells under hypoxia revealed extensive proteoform-level regulation hidden by standard protein-level analysis.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 98%
- Searching for Sulfotyrosines (sY) in a HA(pY)STACK 97%
- The Personalized Proteome: Comparing Proteogenomics and Open Variant Search Approaches for Single Amino Acid Variant Detection 97%
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 98%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 98%
- IceR improves proteome coverage and data completeness in global and single-cell proteomics 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.