Development of Machine Learning Model for Pan-cancer Subgroup Identification using Multi-omics Data
Khadirnaikar, S. R.; Shukla, S.; Prasanna, S. R. M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCancer is a heterogeneous disease and patients with tumors from different organs can share similar epigenetic and genetic alterations. Therefore, it is crucial to identify the novel subgroup of patients with similar molecular characteristics. It is possible to propose a better treatment strategy when the heterogeneity of the patient is accounted for during subgroup identification irrespective of the tissue of origin. In this work, mRNA, miRNA, DNA methylation, and protein expression features from pan-cancer samples were concatenated and non-linearly projected to lower dimension using machine learning (ML) algorithm. This data was then clustered to identify multi-omics based novel subgroups. The clinical characterization of these ML subgroups indicated significant differences in overall survival (OS) and disease free survival (DFS) (p-value<0.0001). The subgroups formed by the patients from different tumors shared similar molecular alterations in terms of immune microenvironment, mutation profile, and enriched pathways. Further, decision-level and feature-level fused classification models were built to identify the novel subgroups for unseen samples. Additionally, the classification models were used to obtain the class labels for the validation samples and the molecular characteristics were verified. To summarize, this work identified novel ML subgroups using multi-omics data and showed that patients with different tumor types can be similar molecularly. We also proposed and validated the classification models for subgroup identification. The proposed classification models can be used to identify the novel multi-omics subgroups and the molecular characteristics of each subgroup can be used to design appropriate treatment regimen.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The repertoire of copy number alteration signatures in human cancer 95%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 95%
- CCLA: an accurate method and web server for cancer cell line authentication using gene expression profiles 94%
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 95%
- A Novel Tool for Multi-Omics Network Integration and Visualization: A Study of Glioma Heterogeneity 94%
- A Comprehensive Targeted Panel of 295 Genes: Unveiling Key Disease Initiating and Transformative Biomarkers in MultipleMyeloma 92%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
- eDAVE - extension of GDC Data Analysis, Visualization, and Exploration Tools 92%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 92%
Similar papers in this journal
- Integrating ensemble systems biology feature selection and bimodal deep neural network for breast cancer prognosis prediction 95%
- Multi-omic signatures identify pan-cancer classes of tumors beyond tissue of origin. 95%
- Identification of miRNA signatures for kidney renal clear cell carcinoma using the tensor-decomposition method 94%
Similar papers in this journal
- Identification of Platform-Independent Diagnostic Biomarker Panel for Hepatocellular Carcinoma using Large-scale Transcriptomics Data 94%
- SCSA: a cell type annotation tool for single-cell RNA-seq data 93%
- GMQN: A reference-based method for correcting batch effects as well as probes bias in HumanMethylation BeadChip 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.