Predicting the similarity of two mass spectrometry runs using only MS1 data
Shouaib, A.; Lin, A.
Show abstract
BackgroundTraditionally researchers can compare the similarity between a pair of mass spectrometry-based proteomics samples by comparing the lists of detected peptides that result from database searching or spectral library searching. Unfortunately, this strategy requires having substantial knowledge of the sample and parameterization of the peptide detection step. Therefore, new methods are needed that can rapidly compare proteomics samples against each other without extensive knowledge of the sample. ResultsWe present a set of neural network architectures that predict the proportion of confidently detected peptides in common between two proteomics runs using solely MS1 information as input. Specifically, when compared to several baseline models, we found that the convolutional and siamese neural networks obtained the best performance. In addition, we demonstrate that unsupervised clustering techniques can leverage the predicted output from our method to perform sample-level characterizations. Our methodology allows for the rapid comparison and characterization of proteomics samples sourced from various different acquisition methods, organisms, and instrument types. ConclusionsWe find that machine learning models, using only MS1 information, can be used to predict the similarity between liquid chromatography-tandem mass spectrometry proteomics runs.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 98%
- nf-encyclopedia: A cloud-ready pipeline for chromatogram library data-independent acquisition proteomics workflows 97%
- Features of peptide fragmentation spectra in single cell proteomics 97%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.