Data-Driven Mathematical Approach for Removing Rare Features in Zero-Inflated Datasets
Ortiz-Velez, A. N.; Kelley, S. T.
Show abstract
Sparse feature tables, in which many features are present in very few samples, are common in big biological data (e.g., metagenomics, transcriptomics). Ignoring the problem of zero-inflation can result in biased statistical estimates and decrease power in downstream analyses. Zeros are also a particular issue for compositional data analysis using log-ratios since the log of zero is undefined. Researchers typically deal with zero-inflated data by removing low frequency features, but the thresholds for removal differ markedly between studies with little or no justification. Here, we present CurvCut, a data-driven mathematical approach to zero-inflated feature removal based on curvature analysis of a "ball rolling down a hill", where the hill is a histogram of feature distribution. These histograms typically contain a point of regime change, a discontinuity with a sharp change in the characteristics of the distribution, that can be used as a cutoff point for low frequency feature removal that considers the data-specific nature of the feature distribution. Our results show that CurvCut works well across a variety of biological data types, including ones with both right- and left-skewed feature distributions, and rapidly generates clear visual results allowing researchers to select data-appropriate cutoffs for feature removal.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life 94%
- Natrix: A Snakemake-based workflow for processing, clustering, and taxonomically assigning amplicon sequencing reads 94%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.