usiGrabber: Automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems
Auge, G.; Clausen, M.; Ketterer, K.; Schaefer, J.; Schmitt, N.; Altenburg, T.; Hartmaring, Y.; Raetz, H.; Schlaffner, C. N.; Renard, B. Y.
Show abstract
MotivationAn unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. ResultsWithin 49 hours, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1,200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under two days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AvailabilityAll code is available at https://github.com/usiGrabber/usiGrabber; the data is available at https://zenodo.org/records/18853258.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 96%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 95%
- Trackable and scalable LC-MS metabolomics data processing using asari 94%