Back

Evaluation of statistical approaches for differential metaproteomics

Hinzke, T.; Kunath, B. J.; Blakeley-Ruiz, J. A.; Korenek, A.; Vintila, S.; Wilmes, P.; Kleiner, M.

2025-12-10 molecular biology
10.64898/2025.12.10.693402 bioRxiv
Show abstract

BackgroundMetaproteomics characterizes and compares molecular phenotypes of organisms in communities by comprehensively analyzing their protein expression profiles using statistical methods. However, not all statistical methods are suitable for determining differentially abundant protein groups in metaproteomic analyses. Statistical challenges in metaproteomics include: data sparsity, non-normality, compositionality, and large between-sample variability. These challenges can potentially be addressed with several data processing steps, including imputation, normalization, transformation, and selection of the appropriate statistical tests. The potential combinations of different processing methods create a complex matrix of analysis options and it is currently unclear how these combinations impact the results of statistical tests on metaproteomic data. ResultsTo determine what data processing methods and statistical tests are best for identifying differentially abundant proteins in metaproteomics datasets, we generated a set of thirteen metaproteomic samples with known compositions, known differences, and differing levels of complexity. These defined metaproteomes address the general challenges outlined above, using various scenarios in metaproteomic data analyses. We compared over 110 different statistical analysis combination options, including regression-based tools, general statistics inference, and machine learning techniques. We found that several combinations within the frameworks of limma, edgeR, MaAslin2, custom linear and Bayesian linear models, and random forests all offer suitable evaluation options. ConclusionsWe highlight key recommendations for differential expression analysis in metaproteomics. Our work enables improved assessment of statistical methods for metaproteomics by establishing a framework for testing statistical approaches, including comprehensive raw mass spectrometry data and reproducible benchmarking code.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.