Back

Comparing Machine Learning Architectures for the Prediction of Peptide Collisional Cross Section

Franklin, E.; Röst, H.

2022-03-04 bioinformatics
10.1101/2022.03.01.482566 bioRxiv
Show abstract

1Mass spectrometry is the method of choice in large-scale proteomics studies. One common method is data-independent acquisition (DIA), which allows for high-throughput analysis of biological samples, but also produces complex data. Methods of peptide separation, in addition to retention time, improve data analysis and there has been increasing interest in separating peptides based on collisional cross section (CCS), which is a measure of the size of a peptide. However, existing libraries that are used during data analysis lack CCS measurements, and this data is expensive and time-consuming to acquire. This has led to the desire to predict library values for mass spectrometry analysis. Here we compare three deep learning architectures, LSTM, CNN, and transformer, for the tasks of retention time and collisional cross section prediction. We show that the LSTM and CNN models perform similarly and that the transformer has a lower performance than expected.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.