RSGSA: a Robust and Stable Gene Selection Algorithm
Saha, S.; Soliman, A.; Rajasekaran, S.
Show abstract
Nowadays we are observing an explosion of gene expression data with phenotypes. It enables researchers to efficiently identify genes responsible for certain medical condition as well as classify them for drug target. Like any other phenotype data in medical domain, gene expression data with phenotypes also suffers from being very underdetermined system. In a very large set of features but a very small sample size domains (e.g., DNA microarray, RNA-seq data, GWAS data, etc.), it is often reported that several different spurious feature subsets may yield equally optimal results. This phenomenon is known as instability. Considering these facts, we have developed a very robust and stable supervised gene selection algorithm to select the most discriminating non-spurious set of genes from the gene expression datasets with phenotypes. Stability and robustness is ensured by class and instance levels perturbations, respectively. We have performed rigorous experimental evaluations using 10 real gene expression microarray datasets with phenotypes. It revealed that our algorithm outperforms the state-of-the-art algorithms with respect to stability and classification accuracy. We have also done biological enrichment analysis based on gene ontology-biological processes (GO-BP) terms, disease ontology (DO) terms, and biological pathways.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 97%
- Ranking Cancer Drivers via Betweenness-based Outlier Detection and Random Walks 95%
- Fast and robust imputation for miRNA expression data using constrained least squares 95%
Similar papers in this journal
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 94%
- Studying the history of tumor evolution from single-cell sequencing data by exploring the space of binary matrices 93%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.