Back

Development of Machine Learning Model for Pan-cancer Subgroup Identification using Multi-omics Data

Khadirnaikar, S. R.; Shukla, S.; Prasanna, S. R. M.

2022-09-17 bioinformatics
10.1101/2022.09.15.507989 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWCancer is a heterogeneous disease and patients with tumors from different organs can share similar epigenetic and genetic alterations. Therefore, it is crucial to identify the novel subgroup of patients with similar molecular characteristics. It is possible to propose a better treatment strategy when the heterogeneity of the patient is accounted for during subgroup identification irrespective of the tissue of origin. In this work, mRNA, miRNA, DNA methylation, and protein expression features from pan-cancer samples were concatenated and non-linearly projected to lower dimension using machine learning (ML) algorithm. This data was then clustered to identify multi-omics based novel subgroups. The clinical characterization of these ML subgroups indicated significant differences in overall survival (OS) and disease free survival (DFS) (p-value<0.0001). The subgroups formed by the patients from different tumors shared similar molecular alterations in terms of immune microenvironment, mutation profile, and enriched pathways. Further, decision-level and feature-level fused classification models were built to identify the novel subgroups for unseen samples. Additionally, the classification models were used to obtain the class labels for the validation samples and the molecular characteristics were verified. To summarize, this work identified novel ML subgroups using multi-omics data and showed that patients with different tumor types can be similar molecularly. We also proposed and validated the classification models for subgroup identification. The proposed classification models can be used to identify the novel multi-omics subgroups and the molecular characteristics of each subgroup can be used to design appropriate treatment regimen.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.