Back

A Novel Glycoproteomics Platform for High-Throughput Identification of Disease-Associated Glycoforms

Wen, S.; Gao, Y.; Miao, X.; Deng, J.; Zhou, Y.; Ge, W.; Bo, S.; Zhang, W.; Zhang, R.; Hou, C.; Ma, J.; Jiang, J.; Yang, S.

2026-03-09 bioinformatics
10.64898/2026.03.06.708406 bioRxiv
Show abstract

Glycosylation is a critical post-translational modification, and its aberrant forms are potent disease biomarkers; however, the comprehensive, site-specific identification of all glycosites and glycoforms across an entire proteome remains prohibitively slow and computationally demanding. To address this bottleneck and accelerate biomarker discovery, we introduce the Glycoproteomics Data Analysis Software (GDAS), a novel, high-throughput platform designed to provide confident, proteome-scale identification of disease-specific glycoforms. GDAS streamlines the analysis through a core, multi-step workflow: it initially employs an ultrafast open search (e.g., MSFragger-Glyco) on mass spectrometry data to rapidly screen and statistically reduce the vast proteome database to a manageable subset of significantly regulated glycoproteins, conserving computational resources for subsequent, in-depth, targeted N- and O-glycosylation analysis using specialized tools (e.g., GlycReSoft and O-Pair). Furthermore, a unique Final Analysis Module utilizes an advanced statistical and machine learning pipeline (incorporating Bootstrap/Bayesian methods, XGBoost, and Random Forest) to integrate quantitative results and generate a robust, comprehensive glycosylation score. We demonstrate GDASs power to recognize biologically relevant glycosylation changes in targeted proteins by validating it using published Alzheimers disease data. GDAS can be downloaded from https://github.com/yang-lab/GDAS. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/708406v1_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@189b270org.highwire.dtl.DTLVardef@1221045org.highwire.dtl.DTLVardef@15a3f50org.highwire.dtl.DTLVardef@1f2d43f_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.