Use of relevancy and complementary information for discriminatory gene selection from high-dimensional cancer data
Haque, M. N.; Sharmin, S.; Ali, A. A.; Sajib, A. A.; Shoyaib, M.
Show abstract
With the advent of high-throughput technologies, life sciences are generating a huge amount of biomolecular data. Global gene expression profiles provide a snapshot of all the genes that are transcribed or not in a cell or in a tissue at a particular moment under a particular condition. The high-dimensionality of such gene expression data (i.e., very large number of features/genes analyzed in relatively much less number of samples) makes it difficult to identify the key genes (biomarkers) that are truly and more significantly attributing to a particular phenotype or condition, such as cancer or disease, de novo. With the increase in the number of genes, simple feature selection methods show poor performance for both selecting the effective and informative features and capturing biological information. Addressing these issues, here we propose Mutual information based Gene Selection method (MGS) for selecting informative genes and two ranking methods based on frequency (MGSf) and Random Forest (MGSrf) for ranking the selected genes. We tested our methods on four real gene expression datasets derived from different studies on cancerous and normal samples. Our methods obtained better classification rate with the datasets compared to recently reported methods. Our methods could also detect the key relevant pathways with a causal relationship to the phenotype.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 97%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 97%
- Cardiac disease diagnosis based on GAN in case of missing data 96%
Similar papers in this journal
- Investigate the relevance of major signaling pathways in cancer survival using a biologically meaningful deep learning model 96%
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 95%
- Combining Gene Ontology with Deep Neural Networks to Enhance the Clustering of Single Cell RNA-Seq Data 95%
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 97%
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 96%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.