Repurposing without Reinventing the Wheel - Ensemble Models for Differential Analysis
Chatterjee, S.; Paul, E.; Liu, Z.; Gao, J.; Basak, P.; Bhattacharyya, A.; Banerjee, C.; Mallick, H.
Show abstract
Inspired by ensemble models in machine learning, we propose a general framework for aggregating multiple distinct base models to enhance the power of published differential association analysis (DAA) methods. We demonstrate this approach by augmenting popular DAA models with one or more biologically motivated alternatives. This creates an ensemble that bypasses the challenge of selecting an optimal model and instead combines the strengths of complementary statistical models to achieve superior performance. Our proposed ensemble learning approach is platform-agnostic and can augment any existing DAA method, providing a general and flexible framework for various downstream modeling tasks across domains and data types. We performed extensive benchmarking across both simulated and experimental datasets spanning single-cell gene expression, bulk transcriptomics, and microbiome metagenomics, where the ensemble strategy vastly outperformed non-ensemble methods, identified more differential patterns than the competing methods, and displayed good control of false positive and false discovery rates across diversified scenarios. In addition to highlighting a substantial performance boost for state-of-the-art DAA methods, this work has practical implications for mitigating the so-called reproducibility crisis in omics data science. An open-source R package implementing the ensemble strategy is publicly available at https://github.com/himelmallick/DAssemble.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scBFA: modeling detection patterns to mitigate technical noise in large-scale single cell genomics data 95%
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 95%
- Knowledge-primed neural networks enable biologically interpretable deep learning on single-cell sequencing data 95%
Similar papers in this journal
- Benchmarking Differential Abundance Analysis Methods for Correlated Microbiome Sequencing Data 95%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 94%
- CosGeneGate Selects Multi-functional and Credible Biomarkers for Single-cell Analysis 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.