Detecting significant components of microbiomes by random forest with forward variable selection and phylogenetics
Dang, T.; Kishino, H.
Show abstract
BackgroundRandom forest (RF) captures complex feature patterns that differentiate groups of samples and is rapidly being adopted in microbiome studies. However, a major challenge is the high dimensionality of microbiome datasets. They include thousands of species or molecular functions of particular biological interest. This high dimensionality significantly reduces the power of random forest approaches for identifying true differences. The widely used Boruta algorithm iteratively removes features that are proved by a statistical test to be less relevant than random probes. ResultWe developed a massively parallel forward variable selection algorithm and coupled it with the RF classifier to maximize the predictive performance. The forward variable selection algorithm adds new variable to a set of selected variables as far as the prespecified criterion of predictive power is improved. At each step, the parameters of random forest are optimized. We demonstrated the performance of the proposed approach, which we named RF-FVS, by analyzing two published datasets from large-scale case-control studies: (i) 16S rRNA gene amplicon data for Clostridioides difficile infection (CDI) and (ii) shotgun metagenomics data for human colorectal cancer (CRC). The RF-FVS approach further screened the variables that the Boruta algorithm left and improved the accuracy of the random forest classifier from 81% to 99.01% for CDI and from 75.14% to 90.17% for CRC. ConclusionValid variable selection is essential for the analysis of high-dimensional microbiota data. By adopting the Boruta algorithm for pre-screening of the variables, our proposed RF-FVS approach improves the accuracy of random forest significantly with minimum increase of computational burden. The procedure can be used to identify the functional profiles that differentiate samples between different conditions.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Microbial Risk Score for Capturing Microbial Characteristics, Integrating Multi-omics Data, and Predicting Disease Risk 98%
- Revealing microbial assemblage structure in the human gut microbiome using latent Dirichlet allocation 97%
- PhyloFunc: Phylogeny-informed Functional Distance as a New Ecological Metric for Metaproteomic Data Analysis 97%
Similar papers in this journal
- Association of Body Index with Fecal Microbiome in Children Cohorts with Ethnic-Geographic Factor Interaction: Accurately Using a Bayesian Zero-inflated Negative Binomial Regression Model 97%
- Identification and predictive machine learning models construction of gut microbiota associated with lymph node metastasis in colorectal cancer 96%
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 95%
Similar papers in this journal
- SPARTA: Interpretable functional classification of microbiomes and detection of hidden cumulative effects. 97%
- MiMeNet: Exploring Microbiome-Metabolome Relationships using Neural Networks 96%
- Individualized network analysis reveals a link between the gut microbiome, diet intervention and Gestational Diabetes Mellitus 96%
Similar papers in this journal
- CAIM: Coverage-based Analysis for Identification of Microbiome 98%
- MTD: a unique pipeline for host and meta-transcriptome joint and integrative analyses of RNA-seq data 96%
- VBayesMM: Variational Bayesian neural network to prioritize important relationships of high-dimensional microbiome multiomics data 96%
Similar papers in this journal
- Phylogeny analysis of whole protein-coding genes in metagenomic data detected an environmental gradient for the microbiota 96%
- MinION Sequencing of colorectal cancer tumour microbiomes - a comparison with amplicon-based and RNA-Sequencing 95%
- Gene co-expression network analysis of the human gut commensal bacterium Faecalibacterium prausnitzii based on WGCNA in R-Shiny 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.