Limitations of de novo sequencing in resolving sequence ambiguity
van Puyenbroeck, S.; Beslic, D.; Suomi, T.; Holstein, T.; Muth, T.; Elo, L. L.; Martens, L.; Bouwmeester, R.; Van Den Bossche, T.; Claeys, T.
Show abstract
De novo peptide sequencing enables peptide identification from fragmentation spectra without relying on sequence databases. However, incomplete spectra create ambiguity, making unambiguous identification challenging. Recent deep learning advances have produced numerous de novo models that predict sequences and refine peptide-spectrum matches under such conditions. Yet, their relative strengths, weaknesses, and ability to handle spectrum ambiguity remain unclear. Here, we benchmark eight state-of-the-art models on three publicly available proteomics datasets, comparing performance using established metrics and quantifying inter-model agreement. We assess post-processing approaches, including iterative refinement, rescoring, and reranking, for their ability to improve identification accuracy, and perform an error analysis to identify common mispredictions and their causes. Model performance varied, with considerable overlap of correct identifications. Post-processing yielded no or only modest improvements. Most sequencing errors were model-independent and driven by limited fragment ion coverage, a limitation also observed in database searches with large search spaces.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pairwise Attention: Leveraging Mass Differences to Enhance De Novo Sequencing of Mass Spectra 97%
- Tailor: non-parametric and rapid score calibration method for database search-based peptide identification in shotgun proteomics 97%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 97%
Similar papers in this journal
- Mistle: bringing spectral library predictions to metaproteomics with an efficient search index 96%
- MSModDetector: A Tool for Detecting Mass Shifts and Post-Translational Modifications in Individual Ion Mass Spectrometry Data 96%
- MS2AI: Automated repurposing of public peptide LC-MS data for machine learning applications 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.