Back

Using Natural Language Processing to Learn the Grammar of Glycans

Bojar, D.; Camacho, D. M.; Collins, J. J.

2020-01-11 bioinformatics
10.1101/2020.01.10.902114 bioRxiv
Show abstract

While nucleic acids and proteins receive ample attention, progress on understanding the structural and functional roles of carbohydrates has lagged behind. Here, we develop a language model for glycans, SweetTalk, taking into account glycan connectivity and composition. We use this model to investigate motifs in glycan substructures, classify them according to their O-/N-linkage, and predict their immunogenicity with an accuracy of [~]92%, opening up the potential for rational glycoengineering.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.