Quantifying the ~75-95% of Peptides in DIA-MS Datasets that were not Previously Quantified
Saxena, G.; Fu, Q.; Binek, A.; Van Eyk, J.
Show abstract
We demonstrate an algorithm termed GoldenHaystack (GH) that, compared to the leading DIA-MS algorithm, (a) quantifies and identifies with better FDR accuracy the peptides found in FASTA search spaces ([~]5-25% of analytes in DIA-MS datasets), (b) quantifies the remaining [~]75-95% of analytes that were previously unquantified, and (c) runs [~]40-200x faster (or [~]1-10x faster than the LC-MS). Specifically, without a FASTA or spectral library, GH can deconvolute and accurately quantify chimeric LC-MS spectra. The central idea that enables this claim is: for sufficiently sized projects (e.g., [≥] [~]50 LC-MS files), pairs of peptides that co-elute in one subset of LC-MS files do not exactly co-elute in a different subset of files. GH thus analyzes a project holistically: it uses multi-partite matching to match fragment ions across all samples, separates and regroups the fragment ions into unique analyte signatures, reduces stochastic noise, and then quantifies those unique analyte signatures.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hybrid Quadrupole Mass Filter Radial Ejection Linear Ion Trap and Intelligent Data Acquisition Enable Highly Multiplex Targeted Proteomics 98%
- Data-Driven Optimization of DIA Mass Spectrometry by DO-MS 98%
- Boosting the MS1-only proteomics with machine learning allows 2000 protein identifications in 5-minute proteome analysis 98%
Similar papers in this journal
Similar papers in this journal
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 99%
- Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectrum Alignment For Discovery of Structurally Related Molecules 97%
- IS-PRM-based peptide targeting informed by long-read sequencing for alternative proteome detection 96%
Similar papers in this journal
- MAFFIN: Metabolomics Sample Normalization Using Maximal Density Fold Change with High-Quality Metabolic Features and Corrected Signal Intensities 97%
- MS2AI: Automated repurposing of public peptide LC-MS data for machine learning applications 96%
- LipidMS 3.0: an R-package and a web-based tool for LC-MS/MS data processing and lipid annotation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.