Back

A Benchmarking Study of Feature Screening Approaches Across Omics Classification Settings

VonKaenel, E.; Bramer, L.; Flores, J.; Metz, T.; Nakayasu, E. S.; Webb-Robertson, B.-J.

2026-02-26 bioinformatics
10.64898/2026.02.25.706632 bioRxiv
Show abstract

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime. Author SummaryA common goal when analyzing high throughput omics data is to identify key biomolecules which may be important in complex biological systems. This type of large-scale analysis is critical, as it can significantly reduce the set of total biomolecules which are considered for more detailed, targeted analysis. Additionally, biomolecules which are highly predictive of specific phenotypes are useful markers to guide treatment. For instance, identifying which biomolecules are predictive of type 1 diabetes could improve early diagnosis, which may improve patient prognosis. However, modern instrumentation can often detect thousands or tens of thousands of biomolecules, and studies typically have a disproportionately limited number of samples. Many measured biomolecules, or features, are often noisy and uninformative, and overcoming this noise to identify an informative set of biomolecules can be challenging for machine learning (ML) models. Fortunately, there are many feature selection strategies which can intelligently reduce the number of "uninformative" biomolecules an ML model needs to overcome. In this work, sure screening, which is a class of feature selection methods that retains the true feature set under certain conditions, is evaluated in the context of omics data analyses. The pros and cons of various sure screening methods, including software availability, are summarized, and their use in the larger feature selection literature is contextualized. Additionally, we benchmark performance by applying a suite of sure screening approaches on several real omics datasets generated to understand the progression of type 1 diabetes.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.