Back

Scalable mass-spectrometry-based molecular phylogeny with TreeMS2

Dierckx, M.; Adams, C.; Gauglitz, J. M.; Bittremieux, W.

2026-03-02 bioinformatics
10.64898/2026.02.27.708420 bioRxiv
Show abstract

Molecular phylogeny is a well-established method for inferring evolutionary relationships from DNA and RNA sequences. Here, we extend this concept beyond genetic information by applying phylogeny-like analysis to proteomic and metabolomic mass spectrometry data, capturing relationships based on the realized molecular phenotype. The resulting phenotype-derived trees can be directly compared with conventional genetic-based trees to identify where molecular phenotypes reflect evolutionary history and where they diverge due to functional adaptation, regulation, or environmental influence. To enable this analysis, we introduce TreeMS2, a computational tool that constructs similarity matrices by directly comparing tandem mass spectrometry (MS/MS) spectra between samples. By bypassing spectrum annotation, TreeMS2 enables rapid, unbiased comparisons. Across diverse datasets, TreeMS2 reconstructs biologically meaningful relationships. In proteomics, phenotype-derived trees recapitulate established taxonomy, with deviations pinpointing sample handling errors. In single-cell proteomics our method distinguishes cell types despite sparse and noisy measurements and in metabolomics it resolves major biochemical divisions and fine-scale compositional structure. Together, these results establish TreeMS2 as a scalable, annotation-independent framework for deriving molecular relationships from raw MS/MS data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.