Attention please: modeling global and local context in glycan structure-function relationships
Dai, B.; Mattox, D. E.; Bailey-Kellogg, C.
Show abstract
Glycans are found across the tree of life with remarkable structural diversity enabling critical contributions to diverse biological processes, ranging from facilitating host-pathogen interactions to regulating mitosis & DNA damage repair. While functional motifs within glycan structures are largely responsible for mediating interactions, the contexts in which the motifs are presented can drastically impact these interactions and their downstream effects. Here, we demonstrate the first deep learning method to represent both local and global context in the study of glycan structure-function relationships. Our method, glyBERT, encodes glycans with a branched biochemical language and employs an attention-based deep language model to learn biologically relevant glycan representations focused on the most important components within their global structures. Applying glyBERT to a variety of prediction tasks confirms the value of capturing rich context-dependent patterns in this attention-based model: the same monosaccharides and glycan motifs are represented differently in different contexts and thereby enable improved predictive performance relative to the previous state-of-the-art approaches. Furthermore, glyBERT supports generative exploration of context-dependent glycan structure-function space, moving from one glycan to "nearby" glycans so as to maintain or alter predicted functional properties. In a case study application to altering glycan immunogenicity, this generative process reveals the learned contextual determinants of immunogenicity while yielding both known and novel, realistic glycan structures with altered predicted immunogenicity. In summary, modeling the context dependence of glycan motifs is critical for investigating overall glycan functionality and can enable further exploration of glycan structure-function space to inform new hypotheses and synthetic efforts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comprehensive analysis of lectin-glycan interactions reveals determinants of lectin specificity 97%
- An Integrated Approach to the Characterization of Immune Repertoires Using AIMS: An Automated Immune Molecule Separator 91%
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 91%
Similar papers in this journal
- Simple and practical sialoglycan encoding system reveals vast diversity in nature and identifies a universal sialoglycan-recognizing probe derived from AB5 toxin B subunits 95%
- CarboGrove: a resource of glycan-binding specificities through analyzed glycan-array datasets from all platforms 94%
- Enhancing the interoperability of glycan data flow between ChEBI, PubChem, and GlyGen. 93%
Similar papers in this journal
Similar papers in this journal
- Syntactic Sugars: Crafting a Regular Expression Framework for Glycan Structures 94%
- nanoBERT: A deep learning model for gene agnostic navigation of the nanobody mutational space 90%
- Bridging Worlds: Connecting Glycan Representations with Glycoinformatics via Universal Input and a Canonicalized Nomenclature 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.