Ratios in Disguise, Truths Arise: Glycomics Meets Compositional Data Analysis
Bennett, A. R.; Lundstrom, J.; Chatterjee, S.; Thaysen-Andersen, M.; Bojar, D.
Show abstract
Comparative glycomics data are an instance of compositional data defined by the Aitchison simplex, where measured glycans are parts of a whole, indicated by relative abundances, which are then compared between conditions. Applying traditional statistical analyses to this type of data often results in misleading conclusions, such as spurious "decreases" of glycans between conditions when other structures sharply increase in abundance, or routine false-positive rates of >25% for differential abundance. Our work introduces a compositional data analysis framework, specifically tailored to comparative glycomics, to account for these data dependencies. We employ center log-ratio (CLR) and additive log-ratio (ALR) transformations, augmented with a model incorporating scale uncertainty/information, to introduce the most robust and sensitive glycomics data analysis pipeline. Applied to many publicly available comparative glycomics datasets, we show that this model controls false-positive rates and results in new biological findings. Additionally, we present new modalities to analyze comparative glycomics data with this framework. Alpha- and beta-diversity enable exploration of glycan distributions within and between biological samples, while cross-class glycan correlations shed light on complex and previously undetected interdependencies. These new approaches have revealed deeper insights into glycome variations that are critical to understanding the roles of glycans in health and disease.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 95%
- GproDIA enables data-independent acquisition glycoproteomics with comprehensive statistical control 95%
- SugarQuant: a streamlined pipeline for multiplexed quantitative site-specific N-glycoproteomics 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- PathIntegrate: Multivariate modelling approaches for pathway-based multi-omics data integration 92%
- Identifying the genes impacted by cell proliferation in proteomics and transcriptomics studies 92%
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.