Back

Protocol for Community-created Public MS/MS Reference Library Within the GNPS Infrastructure

Vargas, F.; Weldon, K. C.; Sikora, N.; Wang, M.; Zhang, Z.; Gentry, E. C.; Panitchpakdi, M. W.; Caraballo, M.; Dorrestein, P. C.; Jarmusch, A. K.

2019-10-15 bioinformatics
10.1101/804401 bioRxiv
Show abstract

RationaleA major hurdle in identifying chemicals in mass spectrometry experiments is the availability of MS/MS reference spectra in public databases. Currently, scientists purchase databases or use public databases such as GNPS. The MSMS-Chooser workflow empowers the creation of MS/MS reference spectra directly in the GNPS infrastructure.\n\nMethodsAn MSMS-Chooser sample template was completed with the required information and sequence tables were generated programmatically. Standards in methanol-water (1:1) solution (1 M) were placed into wells individually. An LC-MS/MS system using data-dependent acquisition in positive and negative modes was used. Species that may be generated under typical ESI conditions are chosen. The MS/MS spectra and MSMS-Chooser sample template were subsequently uploaded to MSMS-Chooser in GNPS for automatic MS/MS spectral annotation.\n\nResultsData acquisition quickly and effectively collected MS/MS spectra. MSMS-Chooser was able to accurately annotate 99.2% of the manually validated MS/MS scans that were generated from the chemical standards. The output of MSMS-Chooser includes a table ready for inclusion in the GNPS library (after inspection) as well as the ability to directly launch searches via MASST. Altogether, the data acquisition, processing, and upload to GNPS took ~2 hours for our proof-of-concept results.\n\nConclusionsThe MSMS-Chooser workflow enables the rapid data acquisition, analysis, and annotation of chemical standards, and uploads the MS/MS spectra to community-driven GNPS. MSMS-Chooser democratizes the creation of MS/MS reference spectra in GNPS which will improve annotation and strengthen the tools which use the annotation information.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.