CDState: an unsupervised approach to predict malignant cell heterogeneity in tumor bulk RNA-sequencing data
Kraft, A.; Yates, J.; Boeva, V.
Show abstract
Intratumor transcriptional heterogeneity (ITTH), defined by the coexistence of diverse cell states within one tumor, complicates cancer treatment by contributing to variable therapeutic responses. Although single-cell RNA sequencing can resolve this complexity, its cost and technical demands limit its large-scale use. Bulk RNA-seq data provide a scalable alternative, but most deconvolution methods depend on predefined references, restricting their ability to detect novel malignant states. Unsupervised approaches avoid these constraints but are not tailored to capture heterogeneity within the malignant compartment. To address these limitations, we introduce CDState, an unsupervised method for inferring malignant cell subpopulations from bulk RNA-seq data. CDState utilizes non-negative matrix factorization improved with sum-to-one constraint and a cosine similarity-based optimization to deconvolve bulk gene expression into distinct cell state profiles. We demonstrate robustness of CDState on bulkified single-cell RNA-seq datasets from five cancer types, showing that it outperforms existing unsupervised deconvolution methods in the estimation of both cell state proportions and gene expression profiles. Applied to 33 cancer types from The Cancer Genome Atlas, CDState reveals recurrent gene programs, including epithelial-mesenchymal transition, MYC targets, and oxidative phosphorylation, as major contributors to malignant cell ITTH. We further link malignant states to patient clinical features, identifying states associated with poor prognosis. We propose an intratumor heterogeneity index and show its association with patient survival, clinical characteristics, and therapeutic response. Finally, we identify mutations and copy number alterations in genes such as TP53, KRAS, PIK3CA, SOX2, and SATB1 as potential genetic drivers of malignant cell ITTH across cancer types.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 96%
- omnideconv: a unifying framework for using and benchmarking single-cell-informed deconvolution of bulk RNA-seq data 96%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
Similar papers in this journal
Similar papers in this journal
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 96%
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 96%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
Similar papers in this journal
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 96%
- SpaTM: Topic Models for Inferring Spatially Informed Transcriptional Programs 96%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.