Proportionality-based association metrics in count compositional data
McGregor, K.; Okaeme, N.; Khorasaniha, R.; Veniamin, S.; Jovel, J.; Miller, R.; Mahmood, R.; Graham, M.; Bonner, C.; Bernstein, C. N.; Arnold, D. L.; Bar-Or, A.; Hart, J.; Marrie, R. A.; O'Mahony, J.; Yeh, E. A.; Zhao, Y.; Banwell, B.; Waubant, E.; Knox, N.; Van Domselaar, G.; Zhu, F.; Mirza, A. I.; Tremlett, H.; Armstrong, H.
Show abstract
MotivationCompositional data comprise vectors that describe the constituent parts of a whole. Data arising from various -omics platforms such as 16S and RNA-sequencing are compositional in nature. However, correlations between features on raw counts have no meaningful interpretation. Metrics of proportionality were formulated to address this problem. However, there is an inherent bias that arises when calculating these metrics empirically on count-based measures due to variability in read depths. ResultsWe quantify the bias introduced by empirically calculating proportionality-based association metrics in count data. Additionally, we propose a means of estimating these metrics within a logit-normal multinomial model in pursuit of more accurate estimates. The model-based estimates are shown to outperform empirical estimates in simulated data, and are additionally applied to a mouse embryonic stem-cell single-cell sequencing dataset as well as a pediatric-onset multiple sclerosis metagenomic dataset. Availability and ImplementationAn R package is available at https://CRAN.R-project.org/package=countprop. Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Addressing Erroneous Scale Assumptions in Microbe and Gene Set Enrichment Analysis 97%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 96%
- SCRaPL: hierarchical Bayesian modelling of associations in single cell multi-omics data 96%
Similar papers in this journal
- A negative binomial latent factor model for paired microbiome sequencing data 96%
- Benchmarking imputation methods for network inference using a novel method of synthetic scRNA-seq data generation 95%
- Statistical Analysis of Variability in TnSeq Data Across Conditions Using Zero-Inflated Negative Binomial Regression 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.