Using regularized regression and biological covariation to impute missing values in quantitative proteomics
Sonnett, M.; Peshkin, L.; Kirschner, M. W.
Show abstract
Proteomics studies analyzing many samples typically generate datasets with missing values where many protein abundances are only quantified in a subset of assayed conditions. While multiplexing with isobaric tags can address this by combining multiple samples into a single injection, missing values are unavoidable when the sample count exceeds the number of available isobaric tags (currently >35). Such missing data complicates the interpretation of large-scale studies across diverse experimental conditions. Here, we introduce a method to impute missing values of relative protein abundance by leveraging measurements from other proteins in the dataset through regularized regression. Our technique, which is applicable to diverse datasets including different cell lines, animals, or biochemical perturbations, capitalizes on the hitherto overlooked biological covariation among protein abundance changes. Our analysis of eight published proteomics datasets reveals a robust imputation capability, achieving a median R2 of 0.55 to 0.8 between imputed and measured data. We demonstrate a similar imputation efficacy in multiple measurement modalities: TMT, DIA, label free, and TMT phosphoproteomics. When examining regression coefficients that were pivotal for accurate data imputation we found that those often mirror known biology. We propose that previously overlooked biological covariation might lead to the generation of novel hypotheses and ultimately advance our understanding of systems level protein organization.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 96%
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 95%
- An adaptive, continuous-learning framework for clinical decision-making from proteome-wide biofluid data 95%
Similar papers in this journal
Similar papers in this journal
- Turnover and replication analysis by isotope labeling (TRAIL) reveals the influence of tissue context on protein and organelle lifetimes 95%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 94%
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.