Revisiting Feature Selection with Data Complexity for Biomedicine
Dong, T. N.; Winkler, L.; Khosla, M.
Show abstract
The identification of biomarkers or predictive features that are indicative of a specific biological or disease state is a major research topic in biomedical applications. Several feature selection(FS) methods ranging from simple univariate methods to recent deep-learning methods have been proposed to select a minimal set of the most predictive features. However, there still lacks the answer to the question of "which method to use when". In this paper, we study the performance of feature selection methods with respect to the underlying datasets statistics and their data complexity measures. We perform a comparative study of 11 feature selection methods over 27 publicly available datasets evaluated over a range of number of selected features using classification as the downstream task. We take the first step towards understanding the FS methods performance from the viewpoint of data complexity. Specifically, we (empirically) show that as regard to classification, the performance of all studied feature selection methods is highly correlated with the error rate of a nearest neighbor based classifier. We also argue about the non-suitability of studied complexity measures to determine the optimal number of relevant features. While looking closely at several other aspects, we also provide recommendations for choosing a particular FS method for a given dataset.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 95%
- Fast and robust imputation for miRNA expression data using constrained least squares 95%
- A Comparison of Embedding Aggregation Strategies in Drug-Target Interaction Prediction 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 97%
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 95%
- A machine-learning Approach for Stress Detection Using Wearable Sensors in Free-living Environments 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.