Computational Pipeline to Identify Gene signatures that Define Cancer Subtypes
Mittal, E.; Parikh, V.; Kirchgaessner, R.
Show abstract
MotivationThe heterogeneous nature of cancers with multiple subtypes makes them challenging to treat. However, multi-omics data can be used to identify new therapeutic targets and we established a computational strategy to improve data mining. ResultsUsing our approach we identified genes and pathways specific to cancer subtypes that can serve as biomarkers and therapeutic targets. Using a TCGA breast cancer dataset we applied the ExtraTreesClassifier dimensionality reduction along with logistic regression to select a subset of genes for model training. Applying hyperparameter tuning, increased the model accuracy up to 92%. Finally, we identified 20 significant genes using differential expression. These targetable genes are associated with various cellular processes that impact cancer progression. We then applied our approach to a glioma dataset and again identified subtype specific targetable genes. ConclusionOur research indicates a broader applicability of our strategy to identify specific cancer subtypes and targetable pathways for various cancers.
Matching journals
The top 13 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 96%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 95%
- Integrating ensemble systems biology feature selection and bimodal deep neural network for breast cancer prognosis prediction 94%
Similar papers in this journal
- Directed Bayesian Networks established functional differences between breast cancer subtypes 95%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 94%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 93%
Similar papers in this journal
- An algorithm for drug discovery based on deep learning with an example of developing a drug for the treatment of lung cancer 93%
- BC-Predict: Mining of signal biomarkers and multilevel validation of cascade classifier for early-stage breast cancer subtyping and prognosis 93%
- Predicting GD2 expression across cancer types by the integration of pathway topology and transcriptome data 92%
Similar papers in this journal
- Identification of Platform-Independent Diagnostic Biomarker Panel for Hepatocellular Carcinoma using Large-scale Transcriptomics Data 94%
- Analysis of Pan-Omics Data in Human Interactome Network (APODHIN) 92%
- Computing Skin Cutaneous Melanoma Outcome from the HLA-alleles and Clinical Characteristics 92%
Similar papers in this journal
- Deep Learning Approach to Identifying Breast Cancer Subtypes Using High-Dimensional Genomic Data 94%
- MOVICS: an R package for multi-omics integration and visualization in cancer subtyping 93%
- Multi-layered network-based pathway activity inference using directed random walks: application to predicting clinical outcomes in urologic cancer 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.