Back

A Statistical Framework for Data Purification with Application to Microbiome Data Analysis

Chung, D.; Ma, Q.; Sun, Z.; Zhao, J.; Liu, Z.

2021-09-13 bioinformatics
10.1101/2021.09.13.460157 bioRxiv
Show abstract

Identification of disease-associated microbial species is of great biological and clinical interest. However, this investigation still remains challenges due to heterogeneity in microbial composition between individuals, data quality issues, and complex relationships among species. In this paper, we propose a novel data purification algorithm that allows elimination of noise observations, which leads to increased statistical power to detect disease-associated microbial species. We illustrate the proposed algorithm using the metagenomic data generated from colorectal cancer patients.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.