CancerSubminer: an integrated framework for cancer subtyping using supervised and unsupervised learning on DNA methylation profiles
Choi, J. M.; Zhang, L.
Show abstract
Human cancer is highly heterogeneous, resulting in variable drug resistance and clinical outcomes. This complexity hinders accurate prognosis prediction and the development of targeted therapies. Molecular subtyping addresses these challenges by grouping cancers into more homogeneous subsets based on molecular characteristics, enabling subtype-specific treatment strategies. Subtyping is crucial for early diagnosis, personalized therapy, and improved survival by capturing differential therapeutic responses. Existing approaches to cancer subtyping fall into supervised and unsupervised categories. Supervised methods, often trained on The Cancer Genome Atlas (TCGA), rely on predefined subtype annotations but face limitations in generalizability and novel subtype discovery. Unsupervised methods, while capable of identifying new subtypes, may overlook widely recognized ones, hindering consistency with established classifications. Multi-omics approaches improve accuracy but are constrained by costs and data collection. We propose CancerSubminer, a hybrid subtyping framework that integrates supervised and unsupervised learning. A subtype classifier is first trained on labeled data, after which clustering is applied to extracted features, with low-confidence samples reassigned to refine subtype boundaries. Model is retrained with the refined subtypes, and adversarial training corrects batch effects and learns domain-invariant features across labeled TCGA and unlabeled external datasets. A subsequent semi-supervised fine-tuning phase aligns subtypes between datasets and designates low-confidence samples as potential novel candidates. CancerSubminer was evaluated on five cancer types, including breast, bladder, brain, kidney, and thyroid cancers, using TCGA methylation data with annotated subtypes and unlabeled datasets from the Gene Expression Omnibus. The framework outperformed state-of-the-art subtyping models (iClusterPlus, iClusterBayes, NEMO) and clustering methods (Spectral, K-means). Kaplan-Meier survival analysis demonstrated significant prognostic separation (p < 0.05) for all cancers, including thyroid cancer where predefined subtypes showed no significance but CancerSubminer-derived subtypes did. These findings highlight CancerSubminers ability to identify distinct prognostic subtypes, mitigate batch effects, and improve prognostic stratification across heterogeneous datasets. CancerSubminer is publicly available at https://github.com/joungmin-choi/CancerSubminer.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Benchmarking computational methods for multi-omics biomarker discovery in cancer 96%
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 95%
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 95%
Similar papers in this journal
Similar papers in this journal
- Pan-cancer identification of clinically relevant genomic subtypes using outcome-weighted integrative clustering 95%
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 95%
- DeepProg: an ensemble of deep-learning and machine-learning models for prognosis prediction using multi-omics data 94%
Similar papers in this journal
- Features fusion or not: harnessing multiple pathological foundation models using Meta-Encoder for downstream tasks fine-tuning 95%
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 95%
- Teacher-student collaborated multiple instance learning for pan-cancer PDL1 expression prediction from histopathology slides 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.