Comparing Machine Learning Architectures for the Prediction of Peptide Collisional Cross Section
Franklin, E.; Röst, H.
Show abstract
1Mass spectrometry is the method of choice in large-scale proteomics studies. One common method is data-independent acquisition (DIA), which allows for high-throughput analysis of biological samples, but also produces complex data. Methods of peptide separation, in addition to retention time, improve data analysis and there has been increasing interest in separating peptides based on collisional cross section (CCS), which is a measure of the size of a peptide. However, existing libraries that are used during data analysis lack CCS measurements, and this data is expensive and time-consuming to acquire. This has led to the desire to predict library values for mass spectrometry analysis. Here we compare three deep learning architectures, LSTM, CNN, and transformer, for the tasks of retention time and collisional cross section prediction. We show that the LSTM and CNN models perform similarly and that the transformer has a lower performance than expected.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Machine learning strategies to tackle data challenges in mass spectrometry-based proteomics 95%
- Automated machine learning and explainable AI (AutoML-XAI) for metabolomics: improving cancer diagnostics 94%
- Deep Learning on Multimodal Chemical and Whole Slide Imaging Data for Predicting Prostate Cancer Directly from Tissue Images 93%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 95%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 94%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.