Quantifying Intratumor Heterogeneity by Key Genes Selected Using Concrete Autoencoder
Tanvir, R. B.; Sobhan, M.; Mamun, A. A.; Mondal, A. M.
Show abstract
The tumor cell population in cancer tissue has distinct molecular characteristics and exhibits different phenotypes, thus, resulting in different subpopulations. This phenomenon is known as Intratumor Heterogeneity (ITH), a major contributor to drug resistance, poor prognosis, etc. Therefore, quantifying the levels of ITH in cancer patients is essential, and many algorithms do so in different ways, using different types of omics data. DEPTH (Deviating gene Expression Profiling Tumor Heterogeneity) is the latest algorithm that uses transcriptomic data to evaluate the ITH score. It shows promising performance, has strong similarity with six other algorithms and has an advantage over two algorithms that uses the same type of data (tITH, sITH). However, it has a major drawback since it uses expression values of all the genes ([~]20K genes) in quantifying ITH levels. We hypothesize that a subset of key genes is sufficient to quantify the ITH level. To prove our hypothesis, we developed a deep learning-based computational framework using unsupervised Concrete Autoencoder (CAE) to select a set of cancer-specific key genes that can be used to evaluate the ITH score. For the experiment, we used gene expression profile data of tumor cohorts of breast, kidney, and lung cancer from the TCGA repository. Using multi-run CAE, we selected three sets of key genes, each set related to breast, kidney, and lung tumor cohorts. For the three cancers stated and three molecular subtypes of lung cancer, we calculated the ITH level using all genes and key genes selected by CAE and performed a side-by-side comparison. We could reach similar conclusions for survival and prognostic outcomes based on ITH scores derived from all genes and the sets of key genes. Additionally, for subtypes of lung cancer, the comparative distribution of ITH scores derived from all and key genes remains similar. Based on these observations, it can be stated that a subset of key genes, instead of all genes, is sufficient for ITH quantification. Our results also showed that many key genes are prognostically significant, which can be used as possible therapeutic targets.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 96%
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 96%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 97%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 97%
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 95%
Similar papers in this journal
- Investigate the relevance of major signaling pathways in cancer survival using a biologically meaningful deep learning model 98%
- Single-Cell Classification Using Graph Convolutional Networks 95%
- Combining Gene Ontology with Deep Neural Networks to Enhance the Clustering of Single Cell RNA-Seq Data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.