DiNovo: high-coverage, high-confidence de novo peptide sequencing using mirror proteases and deep learning
Cao, Z.; Peng, X.; Zhang, D.; Zhou, P.; Kang, L.; Chi, H.; Wu, R.; Cheng, Z.; Zhang, Y.; Dai, J.; Li, Y.; Yao, L.; Li, X.; Yang, J.; Wang, H.; Xu, P.; Fu, Y.
Show abstract
Despite the recent advancements driven by deep learning, de novo peptide sequencing is still constrained by incomplete peptide fragmentation and insufficient protein digestion in current single protease-based proteomic experiments. Here, we present a software system, named DiNovo, for high-coverage and confidence de novo peptide sequencing by leveraging the complementarity of mirror proteases. DiNovo is empowered by several innovative algorithms, including a mirror-spectra recognition algorithm independent of pre-sequencing, two sequencing algorithms based on deep learning and graph theory, respectively, and target-decoy mapping, a method for sequencing result evaluation free of prior peptide identification. Compared with the trypsin protease used alone, DiNovo using two pairs of mirror proteases led to two to three times high-confidence amino acids sequenced. Compared with previous single-protease de novo sequencing algorithms, DiNovo achieved much higher sequence coverages. DiNovo also showed great potential as a powerful complement or alternative to database search for peptide identification with quality control.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 96%
- Mistle: bringing spectral library predictions to metaproteomics with an efficient search index 95%
- Alpha-XIC: a deep neural network for scoring the coelution of peak groups improves peptide identification by data-independent acquisition mass spectrometry 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.