Inferring absolute counts from proportions by constraining multivariate normal distributions
Hage, J.; Koestler, D.; Christensen, B.
Show abstract
Biological measurements often result in proportional data, which derive from underlying biological counts. Proportion data are lacking a dimension of information as compared to counts, restricting available analysis methods and separating the data from the biology. We demonstrate a mathematical technique that estimates absolute counts corresponding to proportion data, which we refer to as Mahalanobis Count Inference (MCI). MCI uses information from a population-representative multivariate normal (MVN) distribution of component counts and ultimately outputs an estimated count and a confidence interval per observation proportion vector. We apply MCI to the imputation of white blood cell (WBC) counts, and of total mRNA within single cells. The method performs very well on total mRNA recapitulation (log-space Pearsons R = 0.81), and well enough on WBC counts to outperform proportions at multiple classification tasks. MCI operates with minimal assumptions, and is applicable to many compositional omics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing 95%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 95%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
Similar papers in this journal
Similar papers in this journal
- Nonlinear ridge regression improves cell-type-specific differential expression analysis 96%
- censcyt: censored covariates in differential abundance analysis in cytometry 95%
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.