A Novel Glycoproteomics Platform for High-Throughput Identification of Disease-Associated Glycoforms
Wen, S.; Gao, Y.; Miao, X.; Deng, J.; Zhou, Y.; Ge, W.; Bo, S.; Zhang, W.; Zhang, R.; Hou, C.; Ma, J.; Jiang, J.; Yang, S.
Show abstract
Glycosylation is a critical post-translational modification, and its aberrant forms are potent disease biomarkers; however, the comprehensive, site-specific identification of all glycosites and glycoforms across an entire proteome remains prohibitively slow and computationally demanding. To address this bottleneck and accelerate biomarker discovery, we introduce the Glycoproteomics Data Analysis Software (GDAS), a novel, high-throughput platform designed to provide confident, proteome-scale identification of disease-specific glycoforms. GDAS streamlines the analysis through a core, multi-step workflow: it initially employs an ultrafast open search (e.g., MSFragger-Glyco) on mass spectrometry data to rapidly screen and statistically reduce the vast proteome database to a manageable subset of significantly regulated glycoproteins, conserving computational resources for subsequent, in-depth, targeted N- and O-glycosylation analysis using specialized tools (e.g., GlycReSoft and O-Pair). Furthermore, a unique Final Analysis Module utilizes an advanced statistical and machine learning pipeline (incorporating Bootstrap/Bayesian methods, XGBoost, and Random Forest) to integrate quantitative results and generate a robust, comprehensive glycosylation score. We demonstrate GDASs power to recognize biologically relevant glycosylation changes in targeted proteins by validating it using published Alzheimers disease data. GDAS can be downloaded from https://github.com/yang-lab/GDAS. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/708406v1_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@189b270org.highwire.dtl.DTLVardef@1221045org.highwire.dtl.DTLVardef@15a3f50org.highwire.dtl.DTLVardef@1f2d43f_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing top-down proteomics of brain tissue with FAIMS 97%
- Detection of Discordant Peptide Quantities in Shotgun Proteomics Data by Peptide Correlation Analysis (PeCorA) 96%
- Boosting the MS1-only proteomics with machine learning allows 2000 protein identifications in 5-minute proteome analysis 96%
Similar papers in this journal
Similar papers in this journal
- Development of a PNGase Rc column for online deglycosylation of complex glycoproteins during HDX-MS 96%
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 96%
- IS-PRM-based peptide targeting informed by long-read sequencing for alternative proteome detection 95%
Similar papers in this journal
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 95%
- An economic and robust TMT labeling approach for high throughput proteomic and metaproteomic analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.