Back

ProteoForge: An Imputation-Aware Framework for Differential Proteoform Discovery in Bottom-Up Proteomics

Ergin, E. K.; Conrrero, A.; Ferguson, K. M.; Lange, P. F.

2025-12-16 bioinformatics
10.64898/2025.12.12.694008 bioRxiv
Show abstract

The human genome contains approximately 20,000 protein-coding genes. However, millions of diverse protein variants, called proteoforms, exist. Despite originating from the same gene, proteoforms often have distinct biological roles. In bottom-up proteomics, the aggregation of peptide measurements into protein-level quantities often obscures this information. Existing methods for proteoform deconvolution are limited by their handling of missing data, which can introduce significant bias. To address this we developed ProteoForge, which builds on an imputation-aware statistical model to identify and group co-varying peptides into quantitatively differential proteoforms (dPFs). Benchmarking against existing deconvolution methods demonstrated that ProteoForge provides high accuracy and stability in datasets with high rates of missing values, complex experimental designs, or varying signal strengths. Application of ProteoForge to proteomics data from lung cancer cells under hypoxia revealed extensive proteoform-level regulation hidden by standard protein-level analysis.

Published in Journal of Proteome Research (predicted rank #1) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.