Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
Xia, J.; Liu, S.; Zhou, J.; Chen, S.; Li, S. Z.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWAccurate protein identification from mass spectrometry (MS) data is fundamental to unraveling the complex roles of proteins in biological systems, with peptide sequencing being a pivotal step in this process. The two main paradigms for peptide sequencing are database search, which matches experimental spectra with peptide sequences from databases, and de novo sequencing, which infers peptide sequences directly from MS without relying on pre-constructed database. Although database search methods are highly accurate, they are limited by their inability to identify novel, modified, or mutated peptides absent from the database. In contrast, de novo sequencing is adept at discovering novel peptides but often struggles with missing peaks issue, further leading to lower precision. We introduce SearchNovo, a novel framework that synergistically integrates the strengths of database search and de novo sequencing to enhance peptide sequencing. SearchNovo employs an efficient search mechanism to retrieve the most similar peptide spectrum match (PSM) from a database for each query spectrum, followed by a fusion module that utilizes the reference peptide sequence to guide the generation of the target sequence. Furthermore, we observed that dissimilar (noisy) reference peptides negatively affect model performance. To mitigate this, we constructed pseudo reference PSMs to minimize their impact. Comprehensive evaluations on multiple datasets reveal that SearchNovo significantly outperforms state-of-the-art models. Also, analysis indicates that many retrieved spectra contain missing peaks absent in the query spectra, and the retrieved reference peptides often share common fragments with the target peptides. These are key elements in the recipe for SearchNovos success. The code for reproducing the results are available in the supplementary materials.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Mistle: bringing spectral library predictions to metaproteomics with an efficient search index 96%
- 3DMolMS: Prediction of Tandem Mass Spectra from Three Dimensional Molecular Conformations 96%
- MSModDetector: A Tool for Detecting Mass Shifts and Post-Translational Modifications in Individual Ion Mass Spectrometry Data 96%
Similar papers in this journal
- SingleFrag: A deep learning tool for MS/MS fragment and spectral prediction and metabolite annotation 93%
- eNODAL: an experimentally guided nutriomics data clustering method to unravel complex drug-diet interactions 92%
- Deep Learning-based Pseudo-Mass Spectrometry Imaging Analysis for Precision Medicine 91%
Similar papers in this journal
- Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data 94%
- Deep Learning Prediction of Glycopeptide Tandem Mass Spectra Powers Glycoproteomics 93%
- Personalized deep learning of individual immunopeptidomes to identify neoantigens for cancer vaccines 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.