A High-Quality Acetylation Dataset Reveals Modest Data Requirements for Transfer Learning to Identify Little Studied Post-Translational Modifications
Hartmaring, Y.; Wang, S.; Jones, A. R.; Vizcaino, J. A.; Schlaffner, C. N.; Renard, B. Y.
Show abstract
Dysregulation of post-translational modifications (PTMs) is associated with severe pathologies, including cancers and Alzheimers disease. Despite their biological importance, identifying modified peptides remains challenging due to the immense combinatorial search space. While searches benefit from prior knowledge of a peptides modification status, the data scarcity for most PTMs hinders the development of accurate deep learning classifiers like AHLF (ad hoc learning of peptide fragmentation). Here, we overcome this data bottle-neck for acetylation and ubiquitination. We harmonised a dataset with about 500,000 high quality acetylated peptide-spectrum matches (PSMs) from nine publicly available acetylation-enriched datasets. We fine-tuned AHLF with the acetylation and a 2-million spectra strong ubiquitination dataset separately and assessed the minimum data requirement for training by iteratively downsampling. Training separate models on SILAC and label-free subsets also assessed the impact of data diversity. The resulting acetylation and ubiquitination models achieve an AUC of 0.87 and 0.90 respectively. Beyond 28,500 acetylated spectra, corresponding to roughly 0.3% of the original models training data, additional data just provides minor performance gains. Finally, we show that data diversity is beneficial for generalizability, while models trained on homogeneous data sources tend to overfit to their respective data type. All code, and model weights are available at https://gitlab.com/dacs-hpi/ahlf-ptmai.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mokapot: Fast and flexible semi-supervised learning for peptide detection 97%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 96%
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 96%
Similar papers in this journal
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 96%
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 95%
- MSModDetector: A Tool for Detecting Mass Shifts and Post-Translational Modifications in Individual Ion Mass Spectrometry Data 94%
Similar papers in this journal
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 96%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 96%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.