Back

msmu: a Python toolkit for modular and traceable LC-MS proteomics data analysis based on MuData

Choi, H.-W.; Lee, B.; Kang, U.-B.; Huh, S.

2026-01-08 bioinformatics
10.64898/2026.01.07.698308 bioRxiv
Show abstract

Computational workflows for MS-based proteomics remain comparatively fragmented, with heterogeneous data formats and analysis pipelines that hinder their reproducibility, interoperability, and reuse of processed data. We present msmu, an open-source Python package that implements a flexible and reproducible end-to-end pipeline for post-search data preprocessing and statistical analysis. At its core, msmu leverages the highly structured MuData format, empowering comprehensive data provenance, transparency in data sharing and reuse, and interoperability with broader Python ecosystem. Together, msmu represents a unique and significant step toward realizing the FAIR (Findable, Accessible, Interoperable, and Reusable) principles in computational proteomics.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.