DLDN-Bench: A Benchmark Framework for Deep Learning de Novo Peptide Sequencing in Proteomics
Schneider, J.; Hartwig, S.; Chadt, A.; Lehr, S.; Al-Hasani, H.; Turewicz, M.
Show abstract
De novo peptide sequencing is an essential approach for analyzing mass spectrometry data because it enables the identification of novel peptides without relying on protein sequence databases. Recent advances in deep learning have substantially improved the performance of de novo sequencing methods, but the rapid emergence of new models has led to heterogeneous evaluation practices and limited comparability. To address this, we introduce DLDN-Bench, a benchmark framework including a set of benchmark datasets derived from human muscle biopsy mass spectrometry data retrieved from PRIDE and annotated through consensus across multiple widely used database search engines. Using these datasets, we systematically benchmark recent deep learning-based de novo sequencing tools alongside traditional approaches. Performance is assessed using established metrics, including precision and coverage relative to a pseudo-ground truth defined by cross-engine agreement. To demonstrate the utility of DLDN-Bench, we benchmark four recent deep learning models and make all results publicly available. This benchmark framework provides a standardized basis for comparing state-of-the-art methods and offers an extensible resource for evaluating future tools in de novo peptide sequencing. Code availabilityhttps://github.com/ddz-icb/DLDN-Bench Data availabilityhttps://doi.org/10.5281/zenodo.19627459
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mistle: bringing spectral library predictions to metaproteomics with an efficient search index 97%
- Alpha-XIC: a deep neural network for scoring the coelution of peak groups improves peptide identification by data-independent acquisition mass spectrometry 97%
- SugarPy facilitates the universal, discovery-driven analysis of intact glycopeptides 96%
Similar papers in this journal
- Improved open modification searching via unified spectral search with predicted libraries and enhanced vector representations in ANN-SoLo 98%
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 96%
- Isobaric matching between runs and novel PSM-level normalization in MaxQuant strongly improve reporter ion-based quantification 96%
Similar papers in this journal
Similar papers in this journal
- promor: a comprehensive R package for label-free proteomics data analysis and predictive modeling 95%
- Covariate balanced allocation of samples to batches to mitigate the impacts of technical variability. 93%
- Enrichment analysis for spatial and single-cell metabolomics accounting for molecular ambiguity 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.