MIPMLP - Microbiome Pre-processing Machine Learning Pipeline.
Jasner, Y.; belogolovski, A.; Ben Izhak, M.; Koren, O.; Louzoun, Y.
Show abstract
16S sequencing results are often used for Machine Learning (ML) tasks. 16S gene sequences are represented as feature counts, which are associated with taxonomic representation. Raw feature counts may not be the optimal representation for ML. We checked multiple preprocessing steps and tested the optimal combination for 16S sequencing-based classification tasks. We computed the contribution of each step to the accuracy as measured by the Area Under Curve (AUC) of the classification. We show that the log of the feature counts is much more informative than the relative counts. We further show that merging features associated with the same taxonomy at a given level, through a dimension reduction step for each group of bacteria improves the AUC. Finally, we show that z-scoring has a very limited effect on the results. These preprocessing steps are integrated into the MIPMLP - Microbiome Preprocessing Machine Learning Pipeline, which is available as a stand alone version at https://github.com/louzounlab/microbiome/tree/master/Preprocess or as a service at http://mip-mlp.math.biu.ac.il/Home ImportanceMicrobiome composition has been proposed as a biomarker (mic-marker) for multiple diseases. However, a clear analysis of the optimal way to represent the gene sequence counts is still lacking. We propose a simple and straight forward method that significantly improves the accuracy of mic-marker studies. This method can be of use to merge two of the most important advances in biology in the last decade: Microbiome analysis, and the introduction of machine learning methods to biological studies.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 95%
- A hybrid CNN-Random Forest algorithm for bacterial spore segmentation and classification in TEM images 94%
- Detecting SARS-CoV-2 lineages and mutational load in municipal wastewater; a use-case in the metropolitan area of Thessaloniki, Greece 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.