Bridging Worlds: Connecting Glycan Representations with Glycoinformatics via Universal Input and a Canonicalized Nomenclature
Urban, J.; Joeres, R.; Bojar, D.
Show abstract
As the field of glycobiology has developed, so too have different glycan nomenclature systems, reflecting diverse cognitive and practical needs of different scientific uses. While each system serves specific purposes, this multiplicity creates challenges for usability, data integration, and knowledge sharing. Here, we present a practical framework for automated nomenclature conversion, taking any nomenclature as input, without having to declare the specific language, and using a canonicalized IUPAC-condensed format as a standardized output representation. Our implementation handles (i) all common nomenclatures, including common typos, (ii) complex cases including structural ambiguities, modifications, and uncertainty in linkage information, and (iii) different compositional representations. This Universal Input framework can translate more than 10 nomenclatures in less than 1 ms, tested on over 50,000 sequences with 95-100% coverage, enabling seamless integration of existing glycan databases and tools while maintaining the specific advantages of each representation system.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comprehensive analysis of lectin-glycan interactions reveals determinants of lectin specificity 93%
- An Integrated Approach to the Characterization of Immune Repertoires Using AIMS: An Automated Immune Molecule Separator 91%
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.