Evaluation of machine learning models for proteoform retention and migration time prediction in top-down mass spectrometry
Chen, W.; McCool, E. N.; Sun, L.; Zang, Y.; Xia, N.; Liu, X.
Show abstract
Reversed-phase liquid chromatography (RPLC) and capillary zone electrophoresis (CZE) are two popular proteoform separation methods in mass spectrometry (MS)-based top-down proteomics. The prediction of proteoform retention time in RPLC and migration time in CZE provides additional information that can increase the accuracy of proteoform identification and quantification. Whereas existing methods for retention and migration time prediction are mainly focused on peptides in bottom-up MS, there is still a lack of methods for the problem in top-down MS. We systematically evaluated 6 models for proteoform retention and/or migration time prediction in top-down MS and showed that the Prosit model achieved a high accuracy (R2 > 0.91) for proteoform retention time prediction and that the Prosit model and a fully connected neural network model obtained a high accuracy (R2 > 0.94) for proteoform migration time prediction.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 98%
- Development of a PNGase Rc column for online deglycosylation of complex glycoproteins during HDX-MS 97%
- IS-PRM-based peptide targeting informed by long-read sequencing for alternative proteome detection 97%
Similar papers in this journal
- MSFragger-DDA+ Enhances Peptide Identification Sensitivity with Full Isolation Window Search 98%
- DeepRTAlign: toward accurate retention time alignment for large cohort mass spectrometry data analysis 97%
- Single Cell Proteomics Using a Trapped Ion Mobility Time-of-Flight Mass Spectrometer Provides Insight into the Post-translational Modification Landscape of Individual Human Cells 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.