Back

Protein Language Model-Aligned Spectra Embeddings for De Novo Peptide Sequencing

NaderiAlizadeh, N.; Dallago, C.; Soderblom, E. J.; Soderling, S. H.

2025-10-03 bioinformatics
10.1101/2025.10.01.679857 bioRxiv
Show abstract

We consider the problem of de novo peptide sequencing in tandem mass spectrometry, where the goal is to predict the underlying peptide sequence given a spectrums fragment peaks and precursor information. We present PLMNovo, a constrained learning framework that leverages pre-trained protein language models (PLMs) to guide the training process. In particular, we cast peptide-spectrum matching as a constrained optimization problem that enforces alignment between spectrum and peptide embeddings produced by a spectrum encoder and a PLM, respectively. We use a Lagrangian primal-dual algorithm to train the spectrum encoder and the peptide decoder by solving the proposed constrained learning problem, while optionally fine-tuning the pre-trained PLM. Through numerical experiments on established benchmarks, we demonstrate that PLMNovo outperforms several state-of-the-art deep learning-based de novo sequencing algorithms.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.