Tree-weighting for multi-study ensemble learners
Ramchandran, M.; Patil, P.; Parmigiani, G.
Show abstract
Multi-study learning uses multiple training studies, separately trains classifiers on individual studies, and then forms ensembles with weights rewarding members with better cross-study prediction ability. This article considers novel weighting approaches for constructing tree-based ensemble learners in this setting. Using Random Forests as a single-study learner, we perform a comparison of either weighting each forest to form the ensemble, or extracting the individual trees trained by each Random Forest and weighting them directly. We consider weighting approaches that reward cross-study replicability within the training set. We find that incorporating multiple layers of ensembling in the training process increases the robustness of the resulting predictor. Furthermore, we explore the mechanisms by which the ensembling weights correspond to the internal structure of trees to shed light on the important features in determining the relationship between the Random Forests algorithm and the true outcome model. Finally, we apply our approach to genomic datasets and show that our method improves upon the basic multi-study learning paradigm.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 94%
- Enabling interpretable machine learning for biological data with reliability scores 93%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 93%
Similar papers in this journal
- Analyzing Biomarker Discovery: Estimating the Reproducibility of Biomarker Sets 93%
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 93%
- ACMTF-R: supervised multi-omics data integration uncovering shared and distinct outcome-associated variation 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.