Back

GlyTrait Brings Insights into Functional Glycosylation

Fu, B.; Wang, G.; Li, C.; Li, Y.; Liu, X.; Zhang, Y.; Lu, H.

2024-07-12 bioinformatics
10.1101/2024.07.09.602632 bioRxiv
Show abstract

Glycomics research often grapples with the interpretability and biological relevance of glycomics data. Using glycosylation derived traits are promising methods for more in-depth biological insights, yet no such bioinformatic tool exists for such task. Here, we developed GlyTrait, a Python-based framework designed to enhance glycomics analysis through the innovative calculation and interpretation of derived traits from N-glycome data. GlyTrait automates the derivation of biologically significant traits, shifting focus from mere glycan abundances to functional glycan properties such as branching and fucosylation. GlyTrait extends the well-established nomenclatures and definitions of derived traits in the N-glycomics community, allowing for fast exploration and analysis of N-glycome data effortlessly. Furthermore, with the well-designed formula grammar, custom derived traits could be materialized without any knowledge of coding. Besides, a two-step post-filtering process reduces information redundancy, maintaining only the most informative traits. Finally, subsequent statistical and interpretable machine learning analysis provide robust insights into the glycosylation patterns associated with disease states. This comprehensive approach not only improves the statistical power and sensitivity compared to traditional methods, but also enhances the interpretability of glycomics data. GlyTraits efficacy is demonstrated through the re-analysis of published glycoengineered CHO cell lines and visceral leishmaniasis patient data, alongside a newly conducted pilot study for hepatocellular carcinoma (HCC) N-glycan biomarker discovery. We are confident in GlyTraits potential to become an indispensable tool for the glycomics community. Significance StatementGlycomics data is often challenging to interpret and analyze for biological relevance. GlyTrait, a Python-based tool, addresses this by automatically calculating biologically significant traits from N-glycome data, focusing on functional properties like glycan branching and fucosylation rather than mere abundance. This tool enhances data interpretability, reduces redundancy through post-filtering, and employs statistical methods to uncover glycosylation patterns linked to diseases. Demonstrated on various published datasets, as well as a newly conducted pilot study with 165 samples for hepatocellular carcinoma N-glycan biomarker discovery, GlyTrait proves to be a powerful tool. It improves the understanding of glycosylation changes in conditions like cancer, thereby benefiting the glycomics community.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.