Beyond P-values: A Multi-Metric Framework for Robust Feature Selection and Predictive Modeling
Chen, R.; Ghosh, A.; Hu, J.; Chen, Y.; Moore, J. H.; Li, R.
10.1101/2025.10.05.680380 bioRxivShow abstract
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does not guarantee predictive utility, and vice versa. Yet few methods unify inferential and predictive evidence within a single selection framework. We introduce MIXER (Multi-metric Integration for eXplanatory and prEdictive Ranking), a domain-agnostic approach that integrates multiple selection metrics into one consensus model via adaptive weighting that quantifies each criterions contribution. Through simulation studies, we demonstrate that different selection metrics identified markedly different feature sets whose over-laps depended on the underlying feature distributions and signal strength. Applied to Alzhemiers disease in UK Biobank, MIXER outperformed every individual criterion, including statistical significance, and generalized to an external disease-specific cohort, Alzheimers Disease Sequencing Project, yielding higher discrimination and stronger risk stratification. The MIXER framwork is also modular and readily extends to other selection criteria and data modalities, providing a practical route to more accurate, interpretable, and transportable predictive models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying and ranking potential driver genes of Alzheimer's Disease using multi-view evidence aggregation 95%
- CLEP: A Hybrid Data- and Knowledge- Driven Framework for Generating Patient Representations 95%
- Distinguishing Specific from Broad Genetic Associations between External Correlates and Common Factors 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 95%
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 94%
- Statistical knockoffs improve biomarker discovery fromtranscriptomic data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.