Empirical-Bayes and Bayesian Hierarchical Modelling forMissingness and Differential Expression in Proteomics
Li, M.; Mallikarjun, V.; Frey, A.; Ogundimu, E.; Trost, M.
Show abstract
Mass spectrometry-based label-free proteomics data often suffer from missing values, especially for low-abundance proteins or when a protein is completely absent in one condition. This makes it challenging to estimate fold changes reliably and perform downstream analyses. Traditional imputation methods often show inconsistent performance across datasets and they typically treat imputed values as fixed rather than uncertain. This can lead to an underestimation of variability in downstream analyses. To address those problems, we present a hierarchical model that accounts for both observed protein intensities and patterns of missing data. Missing values are modelled as left-censored observations below protein-specific detection limits, reflecting the limited sensitivity of the instrument, or being missing with the probability of an intensity dependent manner. Our proposed model captures structure at multiple levels: intensity-level measurements, group-level effects (e.g., experimental conditions), and protein-level variation. To estimate model parameters, we employ an empirical Bayes framework to infer hyperparameters across proteins and use Markov Chain Monte Carlo (MCMC) methods to sample parameters from the posterior distribution. Our pipeline avoids the need for imputing missing values and is designed to produce more reliable fold-change estimates and uncertainty measures. We benchmark our method against existing approaches and demonstrate that it provides more accurate, stable, and robust estimates for differential expression analysis. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=103 SRC="FIGDIR/small/699650v1_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@11da768org.highwire.dtl.DTLVardef@1d9de0aorg.highwire.dtl.DTLVardef@809df1org.highwire.dtl.DTLVardef@16cc3_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 97%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 96%
- VIQoR: a web service for Visually supervised protein Inference and protein Quantification 96%
Similar papers in this journal
- Semi-supervised Bayesian integration of multiple spatial proteomics datasets 96%
- DART-ID increases single-cell proteome coverage 94%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.