Mutually exclusive spectral biclustering and its applications in cancer subtyping
Liu, F.; Yang, Y.; Xu, X. S.; Yuan, M.
Show abstract
Many soft biclustering algorithms have been developed and applied to various biological and biomedical data analyses. However, until now, few mutually exclusive (hard) biclustering algorithms have been proposed although they can be extremely useful for identify disease or molecular subtypes based on genomic or transcriptomic data. We considered the biclustering problem of expression matrices as a bipartite graph partitioning problem and developed a novel biclustering algorithm, MESBC, based on Dhillons spectral method to detect mutually exclusive biclusters. MESBC simultaneously detects relevant features (genes) and corresponding subgroups, and therefore automatically uses the signature features for each subtype to perform the clustering, improving the clustering performance. MESBC could accurately detect the pre-specified biclusters in simulations, and the identified biclusters were highly consistent with the true labels. Particularly, in setting with high noise, MESBC outperformed existing NMF and Dhillons method and provided markedly better accuracy. Analysis of two TCGA datasets (LUAD and BRAC cohorts) revealed that MESBC provided similar or more accurate prognostication (i.e., smaller p value) for overall survival in patients with breast and lung cancer, respectively, compared to the existing, gold-standard subtypes for breast (PAM50) and lung cancer (integrative clustering). In the TCGA lung cancer patients, MESBC detected two clinically relevant, rare subtypes that other biclustering or integrative clustering algorithms could not detect. These findings validated our hypothesis that MESBC could improve molecular subtyping in cancer patients and potentially facilitate better individual patient management, risk stratification, patient selection, therapeutic assignments, as well as better understanding gene signatures and molecular pathways for development of novel therapeutic agents.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A novel single-cell based method for breast cancer prognosis 97%
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 96%
- scPADGRN: A preconditioned ADMM approach for reconstructing dynamic gene regulatory network using single-cell RNA sequencing data 96%
Similar papers in this journal
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 96%
- Differentially Expressed Heterogeneous Overdispersion Genes Testing for Count Data 96%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 96%
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 96%
- VICTOR: A visual analytics web application for comparing cluster sets 96%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.