Back

Beyond P-values: A Multi-Metric Framework for Robust Feature Selection and Predictive Modeling

Chen, R.; Ghosh, A.; Hu, J.; Chen, Y.; Moore, J. H.; Li, R.

2025-12-17 bioinformatics Community evaluation
10.1101/2025.10.05.680380 bioRxiv
Show abstract

High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does not guarantee predictive utility, and vice versa. Yet few methods unify inferential and predictive evidence within a single selection framework. We introduce MIXER (Multi-metric Integration for eXplanatory and prEdictive Ranking), a domain-agnostic approach that integrates multiple selection metrics into one consensus model via adaptive weighting that quantifies each criterions contribution. Through simulation studies, we demonstrate that different selection metrics identified markedly different feature sets whose over-laps depended on the underlying feature distributions and signal strength. Applied to Alzhemiers disease in UK Biobank, MIXER outperformed every individual criterion, including statistical significance, and generalized to an external disease-specific cohort, Alzheimers Disease Sequencing Project, yielding higher discrimination and stronger risk stratification. The MIXER framwork is also modular and readily extends to other selection criteria and data modalities, providing a practical route to more accurate, interpretable, and transportable predictive models.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.