proDA: Probabilistic Dropout Analysis for Identifying Differentially Abundant Proteins in Label-Free Mass Spectrometry
Ahlmann-Eltze, C.; Anders, S.
Show abstract
Protein mass spectrometry with label-free quantification (LFQ) is widely used for quantitative proteomics studies. Nevertheless, well-principled statistical inference procedures are still lacking, and most practitioners adopt methods from transcriptomics. These, however, cannot properly treat the principal complication of label-free proteomics, namely many non-randomly missing values. We present proDA, a method to perform statistical tests for differential abundance of proteins. It models missing values in an intensity-dependent probabilistic manner. proDA is based on linear models and thus suitable for complex experimental designs, and boosts statistical power for small sample sizes by using variance moderation. We show that the currently widely used methods based on ad hoc imputation schemes can report excessive false positives, and that proDA not only overcomes this serious issue but also offers high sensitivity. Thus, proDA fills a crucial gap in the toolbox of quantitative proteomics.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- VIQoR: a web service for Visually supervised protein Inference and protein Quantification 96%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 96%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 96%
Similar papers in this journal
Similar papers in this journal
- Fast alignment of mass spectra in large proteomics datasets, capturing dissimilarities arising from multiple complex modifications of peptides 95%
- Using flux theory in dynamic omics data sets to identify differentially changing signals using DPoP 93%
- MassComp, a lossless compressor for mass spectrometry data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.