Back

LeGenD: determining N-glycoprofiles using an explainable AI-leveraged model with lectin profiling

Li, H.; Peralta, A. G.; Schoffelen, S.; Hansen, A. H.; Arnsdorf, J.; Schinn, S.-M.; Skidmore, J.; Choudhury, B.; Paulchakrabarti, M.; Voldborg, B. G.; Chiang, A. W. T.; Lewis, N. E.

2024-03-30 bioengineering
10.1101/2024.03.27.587044 bioRxiv
Show abstract

Glycosylation affects many vital functions of organisms. Therefore, its surveillance is critical from basic science to biotechnology, including biopharmaceutical development and clinical diagnostics. However, conventional glycan structure analysis faces challenges with throughput and cost. Lectins offer an alternative approach for analyzing glycans, but they only provide glycan epitopes and not full glycan structure information. To overcome these limitations, we developed LeGenD, a lectin and AI-based approach to predict N-glycan structures and determine their relative abundance in purified proteins based on lectin-binding patterns. We trained the LeGenD model using 309 glycoprofiles from 10 recombinant proteins, produced in 30 glycoengineered CHO cell lines. Our approach accurately reconstructed experimentally-measured N-glycoprofiles of bovine Fetuin B and IgG from human sera. Explanatory AI analysis with SHapley Additive exPlanations (SHAP) helped identify the critical lectins for glycoprofile predictions. Our LeGenD approach thus presents an alternative approach for N-glycan analysis. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=136 SRC="FIGDIR/small/587044v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@11687e4org.highwire.dtl.DTLVardef@33b146org.highwire.dtl.DTLVardef@1bb6bcborg.highwire.dtl.DTLVardef@1a1e47e_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.