A Benchmarking Study of Feature Screening Approaches Across Omics Classification Settings
VonKaenel, E.; Bramer, L.; Flores, J.; Metz, T.; Nakayasu, E. S.; Webb-Robertson, B.-J.
Show abstract
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime. Author SummaryA common goal when analyzing high throughput omics data is to identify key biomolecules which may be important in complex biological systems. This type of large-scale analysis is critical, as it can significantly reduce the set of total biomolecules which are considered for more detailed, targeted analysis. Additionally, biomolecules which are highly predictive of specific phenotypes are useful markers to guide treatment. For instance, identifying which biomolecules are predictive of type 1 diabetes could improve early diagnosis, which may improve patient prognosis. However, modern instrumentation can often detect thousands or tens of thousands of biomolecules, and studies typically have a disproportionately limited number of samples. Many measured biomolecules, or features, are often noisy and uninformative, and overcoming this noise to identify an informative set of biomolecules can be challenging for machine learning (ML) models. Fortunately, there are many feature selection strategies which can intelligently reduce the number of "uninformative" biomolecules an ML model needs to overcome. In this work, sure screening, which is a class of feature selection methods that retains the true feature set under certain conditions, is evaluated in the context of omics data analyses. The pros and cons of various sure screening methods, including software availability, are summarized, and their use in the larger feature selection literature is contextualized. Additionally, we benchmark performance by applying a suite of sure screening approaches on several real omics datasets generated to understand the progression of type 1 diabetes.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 94%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 93%
- Correspondence analysis for dimension reduction, batch integration, and visualization of single-cell RNA-seq data 93%
Similar papers in this journal
Similar papers in this journal
- Target-Decoy MineR for determining the biological relevance of variables in noisy data sets 95%
- Soft Windowing Application to Improve Analysis of High-throughput Phenotyping Data 94%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 94%
Similar papers in this journal
- Single sample pathway analysis in metabolomics: performance evaluation and application 95%
- Probabilistic quotient's work and pharmacokinetics' contribution: countering size effect in metabolic time series measurements 94%
- Using flux theory in dynamic omics data sets to identify differentially changing signals using DPoP 93%
Similar papers in this journal
- DIMA: Data-driven selection of a suitable imputation algorithm 93%
- prolfquapp - A User-Friendly Command-Line Tool Simplifying Differential Expression Analysis in Quantitative Proteomics 93%
- Biological Function Assignment Across Taxonomic Levels in Mass-Spectrometry-Based Metaproteomics via a Modified Expectation Maximization Algorithm 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.