A transformer model for de novo sequencing of data-independent acquisition mass spectrometry data
Sanders, J.; Wen, B.; Rudnick, P.; Johnson, R.; Wu, C. C.; Oh, S.; MacCoss, M. J.; Noble, W. S.
Show abstract
A core computational challenge in the analysis of mass spectrometry data is the de novo sequencing problem, in which the generating amino acid sequence is inferred directly from an observed fragmentation spectrum without the use of a sequence database. Recently, deep learning models have made significant advances in de novo sequencing by learning from massive datasets of high-confidence labeled mass spectra. However, these methods are primarily designed for data-dependent acquisition (DDA) experiments. Over the past decade, the field of mass spectrometry has been moving toward using data-independent acquisition (DIA) protocols for the analysis of complex proteomic samples due to their superior specificity and reproducibility. Hence, we present a new de novo sequencing model called Cascadia, which uses a transformer architecture to handle the more complex data generated by DIA protocols. In comparisons with existing approaches for de novo sequencing of DIA data, Cascadia achieves state-of-the-art performance across a range of instruments and experimental protocols. Additionally, we demonstrate Cascadias ability to accurately discover de novo coding variants and peptides from the variable region of antibodies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accounting for digestion enzyme bias in Casanovo 97%
- nf-encyclopedia: A cloud-ready pipeline for chromatogram library data-independent acquisition proteomics workflows 96%
- Extremely fast and accurate open modification spectral library searching of high-resolution mass spectra using feature hashing and graphics processing units 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.